Tenstorrent is turning up the heat in the AI hardware race, and it isn’t being subtle about it. During its TT-Deploy livestream, the company made an attention-grabbing promise: it plans to “crush everybody at everything,” including AI, with its new Galaxy server platform. Big talk aside, Tenstorrent backed the message with performance demos, detailed specs, and pricing for a system designed to take on today’s dominant AI infrastructure.
At the center of Tenstorrent’s new push is Galaxy Blackhole, a fully networked AI server built to scale. The company positions it as a “native AI solution” where compute, memory, and networking are designed as one unified system instead of bolted together from separate parts. The goal is straightforward: better performance for modern AI workloads, especially inference, while keeping costs and efficiency in check when deployed at data center scale.
The engine inside these Galaxy servers is Tenstorrent’s Blackhole chip, built on the RISC-V architecture. RISC-V is increasingly viewed as a serious alternative to ARM and x86 because it’s open and flexible, and Tenstorrent is leaning into that advantage. During the livestream, Jim Keller said the A0 silicon is already shipping, though the company is still working through software bugs.
To explain how Blackhole is built for AI, Tenstorrent pointed to its Tensix “tensor core” design. Each Tensix core includes five RISC processors paired with matrix-multiply units, vector units, and local SRAM. Those RISC processors are fully programmable, and each core is connected through a high-bandwidth network-on-chip (NoC). Tenstorrent then scales the design by packing multiple Tensix cores into a single chip.
Tenstorrent’s messaging also focused heavily on efficiency and economics, not just raw speed. The company argued that some competing inference platforms can increase token throughput only by drastically reducing the number of users being served, which can change the real-world cost equation. Tenstorrent claims its Galaxy servers can maintain strong throughput while delivering a much lower token cost—citing $6 versus roughly $30 on competing platforms—leading to a lower total cost of ownership (TCO) for businesses running large-scale AI services.
The livestream included a generative AI video demo meant to showcase how the Galaxy Supercluster performs on multimedia workloads. Tenstorrent claims up to 10x faster GenAI video performance, showing the system generating an 81-frame 720p video in 2.4 seconds. Put another way, it demonstrated the ability to generate a five-second video clip in less than the time it takes to watch it—faster than real time.
Another headline feature was Blitz Mode, which Tenstorrent describes as an optimization path for premium, latency-sensitive inference. In this mode, Tenstorrent showed Galaxy Blackhole servers running DeepSeek R1-0528 671B at up to 350+ tokens per second per user, using 16 Galaxy servers. The company said this performance outpaces competing GPU-based inference systems, while also supporting batch sizes from 8 to 64 and up to 128K context.
Tenstorrent also highlighted prefill performance (often the painful part of long-context usage), showing sub-4-second time-to-first-token on a 100K context using the same general-purpose Galaxy supercluster setup. For organizations building AI products where responsiveness matters—especially with huge prompts and long context windows—these are the kinds of numbers that can make or break user experience.
For buyers looking at real deployment details, Tenstorrent shared pricing and configuration options. The Galaxy Blackhole server is slated to ship in an air-cooled rack configuration and includes next-generation Blackhole chips plus a fully open-source software stack. Pricing starts at $110,000.
On paper, Tenstorrent says a single system delivers:
23 PFLOPs of FP8 AI compute via 32 Blackhole chips
6.2 GB of on-chip SRAM with 2.9 PB/s bandwidth
1 TB of DRAM with 16 TB/s bandwidth
56 x 800G Ethernet ports for up to 11.2 GB/s of scale-out bandwidth
For larger deployments, Tenstorrent also plans supercluster configurations ranging from 4 to 36 Galaxy servers. A base 4-server configuration starts at $440,000, targeting customers who want scale-out AI infrastructure that can serve serious inference loads across many users.
Tenstorrent’s Galaxy Blackhole announcement is clearly designed to signal a new phase for the company: not just an AI chip maker, but a full-stack platform provider aiming to compete in performance, cost, and scalability. The claims are aggressive, but the specs, demos, and pricing show Tenstorrent is serious about earning a place in the data center AI conversation—especially for organizations prioritizing low latency, long-context inference, and the economics of serving AI at scale.






