
IBM has signed a $240 million agreement with Together AI to bring production-ready, open-source artificial intelligence inference to the cloud. The partnership will create a dedicated pool of Nvidia HGX B300 compute inside IBM Cloud, paired with Nvidia Spectrum-X Ethernet networking and Together AI’s software platform. Together AI will operate the environment and use it to offer enterprise customers high-capacity inference services for open-weight models.
Announced as AI spending increasingly moves from model training to continuous inference, the deal is intended to give organizations an alternative to closed, proprietary models. Instead of depending on a single vendor’s model, enterprises can choose from a range of open-source systems and adapt them to their own data, security policies, and latency requirements.
Open models, flexible deployment
Together AI’s platform has become known for supporting some of the most widely used open-weight models, including DeepSeek, Nemotron, MiniMax, Kimi, and GLM. The new service gives developers access to these models in an environment that is collocated with IBM Cloud’s broader set of enterprise services. That makes it easier to integrate AI outputs with existing business applications, data lakes, identity management systems, and compliance controls.
The companies say the service is designed for a workload profile that is very different from the batch training jobs that dominated previous phases of the AI industry. Inference, which involves using a trained model to generate outputs in real time, often requires low latency, high availability, and the ability to scale up quickly under demand. Those requirements are especially acute for agentic AI systems, which can carry out multistep tasks by repeatedly querying models and acting on the results.
Underneath the hood: Nvidia HGX B300 and Spectrum-X
The infrastructure layer of the deal is central to what IBM is trying to achieve. HGX B300 represents one of the most powerful server designs in Nvidia’s portfolio, combining multiple accelerators per node with high-bandwidth memory to handle both large models and large numbers of concurrent requests. These systems are particularly suited to inference workloads because they can deliver responses quickly without requiring the same kind of scale-up configurations used for frontier training.
Spectrum-X Ethernet networking is the connective tissue that links these systems into an efficient cluster. Nvidia pitched Spectrum-X as a way to bring AI-grade performance to Ethernet networks, an important consideration for organizations that do not want to build proprietary networking infrastructure. For Together AI and IBM, the combination is intended to provide predictable latency and throughput so that enterprises do not have to manage the underlying cluster themselves.
IBM also stressed the hybrid nature of the service. Some workloads will run in IBM Cloud, while other parts of an AI pipeline may remain in an enterprise’s own data center or in an edge environment. IBM has been building tools to move AI workloads between on-premises systems and the cloud, and the partnership with Together AI fits into that strategy. Regulated customers can keep training data or inference calls inside compliant zones, while using the cloud for spikes in demand or experimentation.
A broader IBM-Nvidia collaboration
The partnership is not happening in a vacuum. IBM and Nvidia have expanded their overall collaboration around GPU-native data analytics, document processing, and regulated infrastructure deployments. IBM has also made clear that it sees AI as one of the main reasons customers will move additional mission-critical workloads to its cloud. By adding Together AI as a partner, IBM gains an independent software and services layer that can help differentiate its cloud from hyperscale rivals.
Together AI’s CEO, Vipul Ved Prakash, said the IBM partnership would help the company expand into the enterprise segment and make open-source AI a natural choice for companies that need production-grade inference. “Working alongside IBM with Nvidia gives us that foundation,” he said. “This cluster lets us bring production-grade inference to more companies, faster, and it’s a big step in our push to make open-source AI the obvious choice for enterprises.”
The agreement also highlights the rapid consolidation and expansion of the neocloud market. Together AI competes with a growing set of specialized infrastructure providers that sell raw GPU capacity and higher-level AI services. These neocloud companies have carved out a niche by offering access to expensive hardware in ways that are often faster or more flexible than the largest public cloud providers. For IBM, teaming up with a neocloud rather than trying to build every AI capability internally shortens the path to emerging customers.
An AI infrastructure boom
The decision to invest in AI-optimized infrastructure comes at a moment when market analysts see explosive growth ahead. Gartner has forecast that worldwide spending on AI-optimized infrastructure as a service will grow 96% in 2026, reaching $42 billion. The research firm also expects the market to reach $66 billion in 2027, indicating that this is not a temporary bump but a structural change in how companies buy computing.
In that competitive landscape, IBM Cloud faces a long list of rivals, including AWS, Google Cloud, Microsoft Azure, CoreWeave, Lambda, and many other providers rushing to add GPU capacity. The explosive growth in AI-focused cloud spending has made the market crowded, but the sheer expansion of demand means there is room for multiple providers with different strengths. IBM’s strength has historically been in regulated and large-enterprise environments, and the Together AI service is designed to appeal to those existing customers.
Gartner’s analysts also point to a shift in how AI workloads are consuming cloud resources. In 2026, global spending on inference is projected to reach $23.3 billion, surpassing the $19 billion expected to be spent on training. By 2027, 59% of AI-optimized IaaS spending is forecast to support inference. That underscores why IBM and Together AI are focusing on inference rather than positioning the cluster primarily for foundation-model training.
The rise of agentic AI, in which software agents use models to reason, plan, and act across multiple steps, is expected to amplify this trend. Each agent task can trigger multiple inference calls, making inference the dominant compute consumption model for many AI applications. As organizations shift from model building to production deployment, they are also integrating fine-tuned and domain-specific models into customer-facing systems. These systems require continuous execution rather than occasional training runs.
IBM and Together AI are positioning the new service for exactly that reality. With open models and dedicated compute, the companies want to give developers a predictable environment where AI can be tuned, deployed, and scaled without the constraints of a proprietary API or the operational burden of building a GPU cluster from scratch. The move reflects a broader industry conviction that enterprise AI’s most important phase is now the running of models at scale.
Source:Network World News
