Equinix said Tuesday it will offer AI inference across its network of 281 data centers in 77 metropolitan areas, pairing NVIDIA’s reference architecture with a platform that serves more than 200 open models, in a push by the largest data-center landlord to capture a layer of the AI economy beyond renting space and power.
The offering, called the Equinix Inference Exchange, is being built with NVIDIA and Together AI, an inference platform company. Enterprises would deploy models near their own data and customers rather than sending every request to distant cloud regions, with the first services scheduled for the first quarter of 2027.
Raj Mirpuri, NVIDIA’s vice president for global AI cloud and infrastructure ecosystem, said the companies want to turn Equinix’s interconnected footprint into a distributed fabric for AI inference, so that capacity sits close to the businesses and users generating the requests.
Together AI contributes the software layer. Its platform supports more than 200 open models, giving the exchange a catalog that spans general-purpose and specialized systems, and enterprises can move between them as their needs change, the companies said.
The logic of the offering is geographic. Training runs can tolerate being far from users, because the jobs are scheduled and the results are not waiting on a screen. Inference is different: an answer has to travel back to the person or application asking, so distance shows up as latency, and much of the world’s data is governed by rules that bar it from crossing borders at will.
Equinix’s footprint is densest in the cities where enterprises and their customers already cluster, which is where inference requests originate. Serving a model from a nearby facility shortens the round trip and, for regulated industries, keeps the data inside the country or region where it must remain.
The industry has split along clear lines: training is concentrated in giant clusters built where power is cheap, while inference is spreading toward the edges of the network, where users and data sit. The exchange is Equinix’s attempt to own that second half of the market as it grows.
Equinix’s business model gives the exchange something to stand on. The company has long sold interconnection, the private links that let companies’ systems talk to each other inside the same building or across its footprint. An inference exchange extends that franchise: Equinix hosts the hardware, NVIDIA certifies the reference architecture, Together AI supplies the software, and a customer’s applications can reach the model over the same private connections they already use.
The exchange is designed to run alongside the public clouds rather than against them. Many Equinix customers operate across several clouds and treat the company’s campuses as a neutral meeting point, so an inference service in the same building can sit a network hop away from the applications calling it.
For NVIDIA, the arrangement seeds its hardware into enterprise deployments that may never touch the biggest public clouds, an important channel as inference demand grows beyond the handful of companies building giant training clusters.
The timing matches the market. As companies move AI pilots into production, the volume of inference requests is growing faster than demand for new training runs, and analysts said the bottleneck is shifting to where models are served rather than where they are taught.
The customer base for such services is broadening beyond the AI labs that dominated the first wave of spending. Banks, retailers, manufacturers and health systems are testing models against their own data, and most of them do not plan to build GPU clusters of their own; they want inference where their compliance rules and their users already are.
Open models make the exchange practical. Because the catalog is open-weight, a customer can run the same model at Equinix sites in several cities without licensing constraints or dependence on one vendor’s region map, according to the companies.
Equinix is not alone in chasing the inference layer. Cloud providers sell managed inference, chipmakers bundle their own software stacks, and data-center operators across the industry have been adding AI services to the plain leasing business. The exchange is an unusually direct attempt to own the distribution of models rather than just the buildings they run in.
The economics favor Equinix’s timing. Inference is a recurring workload, billed per request or per deployment, and the margins on serving it are higher than on leasing floor space. As a real-estate investment trust, Equinix is judged on the revenue its buildings produce, and AI services are a way to make each rack pay more than a power bill.
Whether enterprises want to run inference at a colocation provider instead of inside the cloud where their other workloads live is the question the exchange will answer when it opens in 2027. Equinix is betting that the future of AI, many models served close to many customers, looks like the interconnection business it already runs.
For Together AI, the deal opens a distribution channel into thousands of enterprises without the cost of building data centers itself. For NVIDIA, it spreads its reference architecture further through the market. The test of the exchange will be whether customers that start running models at Equinix choose to stay there as their inference volumes grow.


