1 min readfrom TechCrunch

AI infrastructure company Cornelis raises $205M to chip away at Nvidia’s dominance

Our take

Cornelis, an AI infrastructure company, has secured $205 million to challenge Nvidia’s established position. Their innovative approach centers on Active Compute Fabric, a network technology designed to significantly reduce the data latency that often stalls GPU processing. This addresses a critical bottleneck in AI workflows, promising substantial performance gains. Explore how Cornelis is tackling this challenge head-on and empowering more efficient AI development – a concept further examined in our article, "Duplicating baseline benchmarks.”
AI infrastructure company Cornelis raises $205M to chip away at Nvidia’s dominance

The recent $205 million funding round for Cornelis, an AI infrastructure company, signals a growing recognition of a critical bottleneck in the current AI landscape: data movement. While the narrative often centers on the raw power of GPUs, a significant portion of GPU time is, as Cornelis highlights, spent waiting for data. This inefficiency impacts performance and increases costs, a reality explored in discussions around optimizing model training, such as those found in "How to automatically find the batch size when using Accelerate with FSDP2? [D]" — a clear indication of the challenges users face in maximizing resource utilization. The company’s Active Compute Fabric aims to address this head-on, representing a shift towards a more holistic approach to AI infrastructure that prioritizes data flow as much as computational power. This isn't about replacing Nvidia; it's about augmenting the existing ecosystem by tackling a specific, pervasive problem. Nvidia’s own CEO, as discussed in "Nvidia CEO Jensen Huang tells Trump ‘we’re not going to let [an AI slowdown] happen’," is acutely aware of the need to maintain momentum in AI development, and improvements in infrastructure efficiency like those Cornelis is pursuing are crucial to that goal.

The significance of this funding and technology lies in its potential to unlock greater value from existing GPU investments. Currently, organizations are often forced to over-provision GPU resources to compensate for data transfer bottlenecks. Cornelis’ approach promises to optimize data delivery, allowing users to achieve higher performance with the same hardware. This is particularly relevant as the demand for increasingly complex AI models continues to surge. The ability to train and deploy these models efficiently directly translates to faster innovation cycles and reduced operational expenses. The concept of optimizing data flow also resonates with ongoing research into model architectures and training techniques, exemplified by discussions on “Duplicating baseline benchmarks [D],” which underscores the importance of consistent and efficient data handling for reproducible and reliable results. A more streamlined data pipeline allows researchers and engineers to focus on model development rather than wrestling with infrastructure limitations.

The rise of companies like Cornelis reflects a broader trend towards specialization within the AI infrastructure space. While Nvidia has established itself as the dominant player in GPU manufacturing, the supporting ecosystem—networking, storage, and data management—is ripe for disruption. This specialization allows companies to focus on specific pain points and develop targeted solutions. The Active Compute Fabric, if successful, could become a crucial component of future AI deployments, particularly in environments where data volume and velocity are significant factors. It’s a move away from a monolithic, hardware-centric view of AI infrastructure towards a more modular and software-defined approach. This also aligns with the increasingly distributed nature of AI workloads, where data and compute resources are often geographically dispersed, further amplifying the need for efficient data transfer solutions.

Looking ahead, the success of Cornelis will depend on its ability to integrate its Active Compute Fabric seamlessly with existing AI frameworks and hardware. Adoption will require demonstrating clear and measurable performance gains compared to traditional networking solutions. The company’s ability to build a strong developer ecosystem and partner with leading cloud providers will also be critical. Ultimately, the question to watch is whether this focus on data infrastructure can truly challenge Nvidia’s dominance, not through direct competition in GPU manufacturing, but by fundamentally changing how AI workloads are executed and optimized, paving the way for a more efficient and scalable future for artificial intelligence.

The company also announced a product called Active Compute Fabric, a network technology that targets the fact that much GPU time is wasted waiting for data to arrive.

Read on the original site

Open the publisher's page for the full experience

View original article