We compare published DGX Spark, RTX and Ryzen AI Halo inference benchmarks, covering prompt processing, token generation and longer contexts. Alternatives include Framework Desktop, Mac Studio and GB10 systems from other manufacturers.
Decode speed measures output tokens generated per second. Higher values mean faster response generation. All six results use GPT-OSS 20B MXFP4 in Ollama at batch one, from the spreadsheet linked by LMSYS’s DGX Spark review.1
Prefill speed measures input tokens processed per second before response generation. It affects how quickly a model processes a prompt, while decode measures the subsequent output phase.
DGX Spark inference benchmarks
RTX 5090 generated tokens 3.4 times faster than Spark
RTX 5090 generated 205.5 tokens/s, compared with 60.9 on DGX Spark, in the LMSYS spreadsheet. RTX PRO 6000 Blackwell reached 215.2 tokens/s and RTX 5080 reached 140.9. Spark exceeded the two Apple configurations in the same dataset, which recorded 52.7 tokens/s on Mac Studio M1 Max and 46.9 on Mac mini M4 Pro.1
The spreadsheet was retrieved on September 17, 2026. It does not give test dates, prompt lengths or exact Ollama versions for these rows. The review body reports a Spark value of 49.7 tokens/s, while the linked sheet reports 60.906. All six bars use the sheet to keep the source consistent. These results do not cover M3 Ultra or the announced M5 Studio models.2
GPT-OSS 120B decode fell 27% at 32,768 tokens of prior context
DGX Spark generated 58.72 tokens/s with GPT-OSS 120B MXFP4 at zero added context depth in the February 2026 llama.cpp results. At 32,768 tokens of prior context, throughput fell to 42.76 tokens/s, a 27% decline. GPT-OSS 20B fell from 83.43 to 61.65 tokens/s over the same range.3
Prior context is the number of tokens already held in the test context. Both models generated output more slowly as that depth increased, so the short-context rate overstates performance later in a long session for these configurations.
The same file reports GPT-OSS 120B prefill at 2,443.91 tokens/s for a 2,048-token prompt without added depth. A separate test with 32 sequences and 4,096 input tokens each reached 262.20 aggregate decode tokens/s. This is the combined output rate across the batch.3
AMD reports a 4–14% Halo advantage in its May tests
AMD reports higher token throughput for Ryzen AI Halo than DGX Spark on four models: 14% for GLM 4.7 Flash-30B-A3B, 12% for Qwen 3.5-122B-A10B, 7% for GPT-OSS 120B and 4% for Qwen 3.6-35B-A3B. The footnote describes three runs, a 100-token context and Spark software available as of May 6, 2026.4
This is a manufacturer comparison using a preproduction Halo with Ryzen AI Max+ 395 and 128 GB of memory. The public page does not provide the absolute rates or enough quantization and runtime detail to reproduce an equivalent test. The percentages cannot be applied to the February llama.cpp scores to calculate Halo tokens/s.4
DGX Spark alternatives
Spark combines NVIDIA’s GB10 chip, 128 GB of coherent unified memory and 273 GB/s memory bandwidth. It targets local AI development using NVIDIA’s software stack. Its advertised model-size limits describe supported scenarios. Actual fit also depends on weight precision, context, runtime buffers and operating-system memory.5
Availability varies by region and configuration.6478
1. Framework Desktop
Framework’s current configurator lists the 128 GB configuration with Ryzen AI Max+ 395 at $3,449 before optional storage and other selections. Its 32 GB option uses Max 385. Memory is not upgradeable after purchase.6
Developer lhl tested GPT-OSS 120B on Framework Desktop using Vulkan and ROCm, two software paths for running inference on the GPU. The October comparison against published Spark results used different llama.cpp builds and tuned AMD microbatch sizes. Those settings limit how closely the results isolate the hardware difference.9
2. AMD Ryzen AI Halo
Ryzen AI Halo packages Max+ 395, 128 GB of unified memory and ROCm support as a developer platform. AMD’s current page offers the product for purchase and use in the United States. Its May comparison used a $3,999 Halo retail price and $4,699 for Spark. Those are dated price inputs, not a verified September checkout comparison.4
AMD also lists a coming 192 GB Halo based on Ryzen AI Max+ PRO 495. The May benchmark covers the 128 GB Max+ 395 configuration. Performance of the announced 192 GB version remains unverified in this comparison.4
3. Apple Mac Studio
Apple announced M5 Max and M5 Ultra Mac Studio on August 25, 2026. US starting prices are $2,499 and $5,499 respectively. First deliveries are scheduled for September 22, while the 512 GB configuration is scheduled for late October. These are starting prices, not quotes for maximum memory.7
Apple specifies up to 128 GB for M5 Max and 512 GB for M5 Ultra, with up to 1.2 TB/s memory bandwidth on Ultra. Its own AI speedup claims compare selected Apple systems and workloads. They do not establish a Spark comparison. No reproducible M5 Studio-versus-Spark result was accepted into this article’s benchmark dataset.7
Existing M3 Ultra measurements describe the previous generation. M5 purchase comparisons require results for the new hardware and its specific memory configuration.
4. ASUS Ascent GX10 and other GB10 systems
GB10 OEM systems offer alternatives to NVIDIA’s enclosure while retaining the same chip platform. StorageReview compared ASUS Ascent GX10 with NVIDIA’s Founders Edition and systems from Dell, Acer and GIGABYTE. Its vLLM testing found ASUS broadly close to the other systems across the tested models, while its thermal testing examined differences between implementations.8
Cooling, storage, support terms and complete-system price distinguish these implementations. Thermal results remain specific to the tested enclosure and load.
5. A discrete RTX workstation
The LMSYS study provides evidence for considering RTX 5090 or RTX PRO 6000 Blackwell when the workload fits their GPU memory. Both generated GPT-OSS 20B faster than Spark in the historical test.2 NVIDIA specifies 32 GB of dedicated GDDR7 memory for RTX 5090.10
A workstation budget must include the host, power supply, cooling and any additional GPU. Adding several cards increases total installed VRAM, but execution still depends on model partitioning and interconnect behavior. Per-card bandwidth should not be presented as a single pooled memory bandwidth figure.
The LLM memory estimate includes weights, KV cache and runtime overhead. CPU offload changes the execution path when a workload exceeds GPU memory.
Benchmark methodology
The article uses published hardware measurements from three sources. Each comparison retains the reported model and workload, with test dates included where supplied. No new hardware or model-quality test was conducted for this article.
The Spark sweep uses 32-token generation tests (tg32) and 2,048-token prompt-processing tests (pp2048). The five prior-context depths run from zero to 32,768 tokens. The 32-sequence result comes from the separate batched benchmark and measures combined output throughput.3
Concurrent serving also requires latency measurements to assess the effect of additional requests. The separate GPU concurrency benchmark covers this tradeoff on data-center GPUs.
Benchmark limitations
The opening chart uses the linked spreadsheet from LMSYS’s October 2025 review. Its Ollama rows omit prompt/output lengths and runtime versions, so a fully controlled comparison cannot be established. The sheet differs from the review body, and its last measurement date is unknown. NVIDIA supplied early access to Spark for the review.2
The February Spark measurements use a different runtime and test date. They cannot establish an updated RTX-versus-Spark ratio without equivalent peer measurements. AMD’s May percentages are vendor-reported, with absolute rates and exact quantizations undisclosed. The three sources do not form a single controlled ranking.
No complete September 2026 test matrix covering Spark, current Apple hardware, Halo and RTX met the comparison requirements in the sources reviewed. Model-quality parity, power efficiency and total ownership cost were not measured in this research.
Local AI hardware selection
RTX 5090 generated GPT-OSS 20B about 3.4 times faster than Spark in the linked LMSYS dataset. Spark’s 128 GB shared-memory configuration supports a different capacity requirement from RTX 5090’s 32 GB of dedicated memory. Model fit and software compatibility determine which benchmark applies to a deployment.
Framework and Halo provide AMD alternatives, while Mac Studio uses Apple’s GPU software stack. GB10 OEM systems retain the Spark platform with different enclosures and support arrangements. For intermittent workloads, cloud GPU rental options add utilization and transfer costs to the local-versus-cloud comparison.
Further reading
- LLM VRAM Calculator for Self-Hosting
- GPU Concurrency Benchmark: H100 vs H200 vs B200 vs MI300X
- Top 70+ Cloud GPU Providers in 2026
Cite this benchmark
Pick the format that matches where you're publishing. Pasting the link version into your CMS preserves the backlink.
@misc{dilmegani2026,
author = {Dilmegani, Cem and Sarı, Ekrem},
title = {{DGX Spark alternatives: RTX, Ryzen AI Halo and Mac Studio}},
year = {2026},
month = sep,
howpublished = {\url{https://aimultiple.com/dgx-spark-alternatives}},
note = {AIMultiple. Retrieved September 17, 2026}
}Results and timestamps of 19 data points. Download the summary data shown in this article's charts and tables as a ZIP file containing 4 CSV files.
Want the granular data behind it? Join Premium
Changelog
3 updatesAdded an AMD Ryzen AI Halo Mini-PC section to the alternatives comparison.
Removed the Head-to-Head comparison (GPT-OSS 120B Model) section.
Reference Links
Cem's work at AIMultiple has been cited by leading global publications including Business Insider, Forbes, Morning Brew, and Washington Post, global firms like Deloitte and HPE, NGOs like World Economic Forum, and supranational organizations like European Commission. [1], [2], [3], [4], [5]
Throughout his career, Cem served as a tech consultant, tech buyer and tech entrepreneur. He advised enterprises on their technology decisions at McKinsey & Company and Altman Solon for more than a decade. He also published a McKinsey report on digitalization.
He led technology strategy and procurement of a telco while reporting to the CEO. He has also led commercial growth of deep tech company Hypatos that reached a 7 digit annual recurring revenue and a 9 digit valuation from 0 within 2 years. Cem's work in Hypatos was covered by leading technology publications like TechCrunch and Business Insider.
Cem regularly speaks at international technology conferences. He graduated from Bogazici University as a computer engineer and holds an MBA from Columbia Business School.
Be the first to comment
Your email address will not be published. All fields are required. Comments are left in their original language.