RTX 5070 Ti Runs a 35B AI Model, But DDR4 ECC RAM Fails Under the Pressure
Running large language models locally is becoming more practical for enthusiasts, developers, and AI power users, especially with modern GPUs offering faster VRAM and stronger compute performance. One recent case, however, shows that pushing bigger AI models beyond GPU memory can place serious stress on system RAM, particularly older DDR4 modules.
A user running an RTX 5070 Ti with 16GB of GDDR7 VRAM had no trouble loading a 20B AI model directly into the GPU’s memory. For many local AI workloads, that setup would be more than enough. But for more demanding database-related tasks, the user felt the 20B model was too limited and decided to move to a larger 35B model, Ornith-1.5-35B-A3B.
Because a 35B LLM requires far more memory than the RTX 5070 Ti could provide on its own, the system relied heavily on 128GB of DDR4 ECC workstation memory. The configuration reportedly used multiple 8GB DDR4 ECC sticks to help offload the model when GPU VRAM was filled.
At first, the setup appeared to work. The user valued the improved output quality of the 35B model and noted that performance remained usable even after the GPU’s 16GB VRAM was exhausted and the workload spilled over into system memory. That is one of the reasons many local AI users experiment with hybrid GPU and RAM setups, especially when trying to run large AI models without investing in much more expensive hardware.
But after only about 90 minutes of running the 35B model, the system crashed.
After troubleshooting, the user discovered that one of the 8GB DDR4 ECC memory sticks had failed. Assuming it was simply a faulty module, the user removed the bad stick and continued running the same AI model.
The problem did not end there. After 8 to 9 days of running the workload, system logs began showing errors on another DDR4 stick, suggesting that a second module was also beginning to fail.
The user argued that thermals were not the cause, pointing out that the CPU and GPU temperatures stayed below 80°C. However, that only tells part of the story. Memory temperature readings were not provided, and RAM can run much hotter than expected when exposed to sustained, heavy workloads, especially in a workstation packed with multiple DIMMs.
Large AI models can create long periods of continuous memory activity. When the GPU runs out of VRAM and starts leaning on system RAM, the memory may be forced into prolonged sequential reads and writes. That kind of workload is very different from typical desktop use, gaming, or short bursts of productivity work.
DDR4 ECC memory is designed for reliability, and ECC can detect and correct certain memory errors, but it does not make RAM immune to wear, heat, age, or electrical stress. If the modules were older, already weakened, or operating with limited airflow, running a large 35B LLM continuously could have accelerated their failure.
This case raises an important point for anyone interested in running large AI models locally: system RAM matters, but not all memory setups are ideal for AI workloads. While DDR4 can still be useful, especially in older workstations with large memory capacities, it may not be the best option for sustained LLM inference when the workload constantly spills beyond VRAM.
DDR5 memory may be better suited for these scenarios thanks to higher bandwidth, newer designs, and improved operating tolerances. Even then, users should pay close attention to cooling, motherboard airflow, memory temperatures, and long-term stability testing. Large AI models can stress a system in ways that many consumer and workstation builds were not originally designed to handle.
The RTX 5070 Ti itself appears to have handled the workload well, and the 16GB of GDDR7 VRAM was enough for the smaller 20B model. The issue emerged when the user stepped up to a denser 35B AI model and depended heavily on system memory to fill the gap.
For AI enthusiasts, the takeaway is clear: running bigger local LLMs is not just about having a powerful GPU. VRAM capacity, system RAM quality, memory cooling, and workload duration all matter. A setup that works for short testing sessions may not survive days of continuous large model inference without proper thermal monitoring and hardware planning.
In this case, the exact cause of the DDR4 ECC failures remains uncertain. Age, heat, sustained memory stress, and insufficient airflow are all possible factors. Without direct memory temperature data, it is impossible to say with certainty what killed the first stick and began degrading the second.
Still, the incident is a useful warning. If you plan to run 35B AI models or larger on a local PC, especially with part of the model offloaded to system RAM, make sure your memory is healthy, well-cooled, and tested under extended load. Local AI can be powerful, but as models get larger, the demands on every part of the system increase dramatically.






