From Data Centers to Token Factories: TuringData Shares Its Vision for AI-Native Infrastructure at Asia Tech x Singapore 2026

← Back to blog

Singapore, May 2026 — As artificial intelligence continues to advance rapidly across industries, the conversation has shifted from simply acquiring more compute to optimizing the entire data stack. This was the central theme at the AI Summit Stage during Asia Tech x Singapore 2026, where TuringData Vice President Nikhil Madan outlined the next evolution of AI infrastructure in the era of agentic intelligence and large-scale token generation.

Reframing the Infrastructure Conversation

A structural shift is underway in the AI industry: data centers are rapidly evolving into “Token Factories,” where massive volumes of AI tokens are continuously generated and consumed by increasingly autonomous AI systems.

As enterprises scale AI workloads—particularly generative AI and agent-based applications—GPU demand continues to grow exponentially. However, a critical and often overlooked constraint remains:The real bottleneck in AI systems is not compute—it is data delivery.

“The perception is that we just need more GPUs. The reality is that compute is starving. GPUs sit idle waiting for data that never arrives fast enough. Data I/O and storage throughput are the true limiting factors in AI cluster performance, yet they remain chronically underinvested,” said Nikhil Madan.

The result is expensive GPU infrastructure operating well below its full potential, stalled AI deployments, and infrastructure investments that fail to deliver the expected return.

The Token Factory Era Demands a New AI Storage Paradigm

During the session, Nikhil highlighted the key challenges enterprise AI teams consistently face:

  • Prohibitive upfront costs. Traditional AI storage architectures often require substantial capital investment before meaningful workloads can be deployed, creating a high barrier to entry.
  • Network bottlenecks. Large-scale GPU clusters generate significant traffic congestion that limits throughput and reduces training efficiency.
  • Stranded legacy data. Many enterprises sit on vast volumes of historical data that remain siloed and difficult to activate, despite their potential value for training and RAG applications.
  • Inference latency. As AI moves from training into production, traditional storage architectures can introduce latency that undermines real-time inference performance.

The solution is no longer simply adding more compute. It requires a fundamentally different approach to the data layer.

Start pragmatically and scale with the workload—not ahead of it. Keep GPUs continuously fed with data, because idle compute is wasted capital and GPU utilization remains one of the clearest measures of infrastructure efficiency. Rethink storage for the inference era: as Jensen Huang noted at GTC 2026, KV Cache requires a fundamentally different storage architecture—one that is high-throughput, ultra-low-latency, and purpose-built for AI inference. And finally, activate what enterprises already own: the historical data locked inside legacy systems remains one of the highest-value assets available for training and retrieval workflows.

TuringData's Full-Stack AI Storage Vision: Pragmatic, High-Performance, and Future-Ready

Nikhil also presented TuringData's vision for an AI-native data infrastructure stack designed to eliminate bottlenecks across the entire AI lifecycle.The TuringData portfolio consists of three core components:

  • TuringData Platform — A high-performance parallel file system built for large-scale AI training and data pipelines, delivering up to 480 GB/s throughput and 7.5 million IOPS from a three-node cluster.
  • TuringData Cache Fabric — A multi-layer KV Cache management solution with native support for vLLM, SGLang, and NVIDIA Dynamo, delivering up to a 91% reduction in Time to First Token (TTFT) and an 8× increase in token throughput.
  • TuringData Flash — NVMe-based storage appliances designed for workloads requiring extreme throughput and ultra-high IOPS.

Together, these solutions form a unified architecture designed not to replace existing infrastructure, but to unlock and accelerate it—allowing organizations to scale with their workloads rather than overbuild in advance.

A key differentiator is accessibility: TuringData's architecture allows enterprises to deploy starting with as few as three nodes. This drastically minimizes day-one CAPEX and enables companies to prove ROI before scaling up.

Looking Ahead

As AI continues to evolve toward agentic systems and token-driven economies, TuringData remains focused on building the foundational data infrastructure required for scalable, cost-efficient, high-performance AI.

With deployments spanning AI research labs, GPU cloud providers, telecommunications operators, financial institutions, physical AI, and autonomous driving, TuringData is helping power the next generation of AI-native infrastructure.

Find out how much GPU utilization you're leaving on the table. Schedule a free 30-minute architecture review with our experts.