A close-up of a XuanTie chip on a purple-lit circuit board.

Alibaba’s 5nm XuanTie C950 Brings Qwen-3.8 27B to Native RISC-V AI Computing

Alibaba is making a major push to control more of its artificial intelligence stack by bringing day-zero support for its Qwen-3.8 27B AI model to its own XuanTie C950 RISC-V chip. The move signals a broader strategy: reduce reliance on outside hardware platforms, optimize AI models for in-house silicon, and build a stronger ecosystem around custom compute.

The Qwen-3.8 27B model is especially notable because it is a 27-billion-parameter open-weight AI model capable of running with relatively modest hardware requirements. It can operate on systems with just 32GB of VRAM, making it far more accessible than many large AI models that demand expensive, high-end accelerator setups. Its coding performance has been described as competitive with leading frontier models, yet it remains efficient enough to run on a single MacBook-class machine.

Now, Alibaba has optimized that same model for its XuanTie C950 processor, a 64-core RISC-V chip designed for edge AI and private inference workloads. According to reported performance figures, the C950 can run the Qwen-3.8 27B model at around 30 tokens per second, with a Time To First Token of roughly 1.9 seconds. For AI inference, those numbers are meaningful because they point to fast response times without depending on a traditional GPU-based setup.

The XuanTie C950 was introduced as a server-grade 64-bit RISC-V processor built for demanding AI workloads. Instead of using GPUs as the main engine for inference, Alibaba integrated matrix and vector acceleration directly into the processor. This allows the chip to handle core AI operations natively, reducing the need for separate graphics processors in certain deployment scenarios.

The chip includes 64 compute cores on a single piece of silicon, with clock speeds scaling up to 3.20GHz. These cores are organized into clusters of eight and connected using high-speed AMBA CHI fabric, helping data move efficiently across the processor. That design is important for AI workloads, where memory movement and inter-core communication can become major performance bottlenecks.

Alibaba also equipped the XuanTie C950 with a standard L1 cache structure, a flexible L2 cache, and optional shared L3 cache support. To further improve efficiency, the chip uses hardware-level intelligent data prefetching, allowing it to predictively load data into cache before the execution engine requests it. In real-world AI inference, that can help reduce delays and improve sustained throughput.

The use of the open-source RISC-V instruction set is another key part of Alibaba’s strategy. By building on RISC-V, Alibaba avoids the licensing costs tied to x86 and Arm-based architectures while gaining more freedom to customize the chip for its own AI workloads. This gives the company more control over hardware design, software optimization, and long-term product direction.

The XuanTie C950 also features an 8-instruction decode width, enabling each core to read and process a large number of instructions at once. Its 16-stage pipeline is designed to balance high clock speeds with the ability to execute complex server and AI tasks efficiently. Together, these architectural choices make the chip better suited for running smaller and mid-sized large language models directly on the processor.

Unlike GPU-based systems built for massive parallel workloads and high-concurrency public AI services, the XuanTie C950 appears better suited for edge deployment, enterprise inference, and private AI workloads. It runs a single inference thread per socket, which may limit its role in large public API environments but makes it attractive for controlled, secure, and localized AI deployments.

One of the most important claims around the C950 is that it can run billion-parameter large language models natively, without emulation or translation layers. That matters because native execution can improve efficiency, reduce overhead, and allow the hardware to execute AI model operations more directly. The integrated acceleration engines and custom instruction support are designed specifically for the kinds of matrix and vector calculations used in modern AI models.

Reports suggest the XuanTie C950 may be manufactured using TSMC’s 5nm process, although Alibaba has not officially confirmed the fabrication details. If accurate, that would place the chip on a modern manufacturing node, helping with power efficiency and performance density.

Alibaba’s decision to bring immediate Qwen-3.8 27B support to the XuanTie C950 is more than a technical update. It is a strategic move. By making its own AI models run efficiently on its own processors, Alibaba can create a tighter hardware-software ecosystem. This mirrors the broader trend in AI computing, where companies increasingly want control over everything from model architecture to silicon design.

The benefits could be significant. Alibaba can deploy the C950 in edge AI products, enterprise systems, and data centers while pairing it with other AI accelerators where needed. This flexibility gives the company more options for handling inference workloads efficiently and cost-effectively.

For customers, the appeal is clear: a powerful open-weight AI model, optimized for custom RISC-V hardware, with fast token generation and low response latency. For Alibaba, the bigger prize is ecosystem control. If developers and enterprises begin building around Qwen models and XuanTie chips, Alibaba gains a stronger position in the rapidly expanding AI infrastructure market.

The arrival of Qwen-3.8 27B on the XuanTie C950 shows how quickly AI hardware is evolving beyond traditional GPU-centric designs. While GPUs will remain essential for many large-scale training and inference tasks, specialized CPUs with integrated AI acceleration could become increasingly important for edge computing, private AI assistants, coding tools, and enterprise inference.

Alibaba’s latest move makes one thing clear: the future of AI will not be defined only by who builds the biggest models, but also by who can run them efficiently, affordably, and at scale on optimized hardware.