Old Lenovo Laptop Runs a 27B Local AI Model After Being Upgraded With a Desktop Radeon RX 7900 XT
Running large local AI models on an older laptop usually sounds impossible, especially when the machine does not have a dedicated graphics card. Large language models often need serious GPU power and plenty of VRAM, which most notebooks simply cannot provide. However, one Lenovo laptop owner found a creative workaround: connect a full-size desktop AMD Radeon RX 7900 XT through the laptop’s M.2 slot.
The result is an unusual but impressive DIY setup that allows the old notebook to run Qwen3.6 27B at surprisingly strong speeds.
Because the Lenovo laptop only had one available M.2 slot, the project required a few compromises. Instead of using the internal slot for storage, the user connected an ADT-Link PCIe riser adapter to the M.2 interface and attached the Radeon RX 7900 XT externally. Since that slot was now occupied by the graphics card connection, Windows had to be booted from an external drive through USB.
The RX 7900 XT also needed its own power source. A laptop cannot supply enough power for a high-end desktop GPU, so the graphics card was connected to a separate 750W power supply. This turned the notebook into a hybrid desktop-laptop AI workstation, with the external GPU handling the heavy workload.
To reduce unnecessary memory usage, the laptop’s built-in display was kept on integrated graphics, while an external monitor was connected directly to the Radeon RX 7900 XT. This helped prevent extra VRAM overhead and allowed more of the GPU’s 20GB of GDDR6 memory to be used for the AI model.
That 20GB VRAM capacity is the key reason the setup worked so well. Large language models such as Qwen3.6 27B can be extremely demanding, but the RX 7900 XT’s large framebuffer allowed the model weights and KV cache to be offloaded to the GPU. Using llama.cpp, the system reportedly reached around 55 to 60 tokens per second, which is a very respectable result for a machine that was never designed for this type of workload.
The biggest limitation was not the GPU, but the laptop’s system memory. With a context length of 100K tokens, the entire 16GB of system RAM was consumed. That makes memory the main bottleneck in this experiment. More RAM would likely improve the overall experience, but upgrading memory is not always easy or affordable, especially with current DRAM pricing challenges.
Even with that limitation, the project is a great example of how older hardware can be pushed far beyond its original purpose. Instead of replacing the entire laptop, the owner used an available M.2 slot, an external PCIe adapter, a desktop GPU, and a separate power supply to create a capable local AI machine.
This kind of setup is not exactly practical for everyday laptop use. It is bulky, requires extra hardware, and depends on a careful configuration. However, for users interested in local LLM performance, AI experimentation, and hardware modding, it shows what is possible with the right parts and a bit of persistence.
The experiment also highlights why VRAM has become so important for local AI workloads. While gaming performance still matters for GPUs, large memory capacity is increasingly valuable for running AI models locally. In this case, the Radeon RX 7900 XT’s 20GB of VRAM allowed an aging Lenovo notebook to handle a demanding 27B-parameter model that would overwhelm many lower-memory graphics cards.
For anyone trying to run local AI models without buying a brand-new workstation, this project is an inspiring reminder: sometimes the most interesting performance upgrades come from creative hardware hacks rather than standard upgrades.






