TuringData Debuts in MLPerf® Storage with 541.5 GiB/s from Three Nodes

← Back to blog

Our first MLPerf® Storage submission supported 96 simulated accelerators in 3D U-Net, demonstrating sustained AI data performance from a lean architecture.


1 September , 2026 

Today, we are announcing TuringData’s first results in the MLCommons MLPerf® benchmark suite.

A three-node TuringData Flash deployment delivered up to 541.5 GiB/s of aggregate bandwidth in the 3D U-Net workload while supporting 96 simulated accelerators.

For our first submission, this result demonstrates an important capability: supplying accelerator-intensive workloads with sustained data performance without requiring a large storage footprint.

Built for Today’s AI Compute Environment

As AI clusters increase in size, storage performance has a direct effect on accelerator utilization. When data cannot reach compute infrastructure quickly enough, expensive accelerators spend more time waiting and the cost of producing useful work increases.

TuringData Flash is designed to address this challenge through a lean, scale-out architecture that keeps data moving across training, checkpointing and inference-related workflows.

The MLPerf® result highlights three core capabilities:

  1. Sustained bandwidth at accelerator scale. TuringData Flash is designed to maintain consistent data delivery as accelerator counts and datasets grow.
  2. High performance from a compact configuration. The three-node deployment demonstrates a practical starting point for organizations building or expanding AI infrastructure.
  3. A unified data path across AI workflows. The same platform supports training, checkpoint, as well as inference-related IO.

Our MLPerf® Storage v3.0 Results

The three-node submission delivered:

  • 3D U-Net: 541.5 GiB/s read throughput, supporting 96 simulated accelerators
  • Checkpoint-70B: 539.9 GiB/s read and 307.1 GiB/s write
  • Checkpoint-8B: 166 GiB/s read and 108.7 GiB/s write
  • KV Cache: 293.9 GiB/s read and 59.4 GiB/s write

*GiB/s = gibibytes per second.

The Architecture Behind

The TuringData Flash comprises all-flash distributed storage appliances powered by the TuringData Platform file system.

The appliances combine NVMe SSDs, 400 Gbps InfiniBand/RoCE or Ethernet network, and support for NVIDIA GPUDirect Storage. The software-defined architecture brings together an elastic data network, distributed metadata and a multi-tenant global namespace etc. enterprise features.

CTA Image

From Benchmark to Production

Explore TuringData Flash

A Proud First Submission

Submitting to MLPerf® Storage for the first time required careful engineering, extensive validation and close collaboration across our team. We are proud of the result, but we see it as a beginning rather than a finish line. We want to contribute meaningfully to the continued development of AI storage and the broader AI infrastructure ecosystem. That means continuing to test our technology against demanding, transparent benchmarks; learning from the wider engineering community; and turning those insights into systems that are faster, more efficient and easier to operate.

As AI models, datasets and clusters continue to grow, storage must evolve alongside them. Our focus is not simply on producing a larger headline number. It is on helping AI infrastructure sustain performance across data ingestion, distributed training, checkpointing, retrieval and other increasingly important stages of the AI data lifecycle. We look forward to sharing more of what we learn and contributing our perspective as the AI storage landscape develops.

Built in Singapore, Contributing Globally

TuringData’s story is rooted in Singapore—a global technology and data-centre hub that connects Asia with the wider world. Building from Singapore has given us a close view of both regional and global infrastructure challenges. Organizations everywhere are increasing their investment in AI while navigating practical considerations such as data sovereignty, infrastructure density, operational efficiency, energy consumption and access to high-performance computing resources.

ASEAN is an important part of our growth. We see significant opportunities to support the region’s neo-cloud providers, data-centre operators, AI infrastructure builders and enterprise AI teams as they move from experimentation to production-scale deployment. But our ambition extends well beyond Southeast Asia. We believe valuable infrastructure innovation can emerge from Singapore and contribute to the global AI ecosystem.

What Comes Next

Our first MLPerf® Storage submission gives us a strong foundation. From here, we will continue refining our architecture, participating in rigorous industry evaluation, expanding internationally and engaging with the broader AI and storage communities. We are grateful to MLCommons for establishing an open and rigorous benchmark, and to everyone at TuringData whose work made this submission possible. While we are proud of what this result represents, we are even more excited about what comes next.


TuringData is building the AI data infrastructure layer for accelerated training and smart inference. Our product portfolio includes the TuringData Platform, TuringData Flash and TuringData Cache Fabric.

Full benchmark results and configuration details are available through the MLCommons Storage benchmark page.

MLCommons® and MLPerf® are registered trademarks of MLCommons Association in the United States and other countries.

Find out how much GPU utilization you're leaving on the table. Schedule a free 30-minute architecture review with our experts.