Index  ›  business  ›  Forbes
business · Forbes ↗

Modular Delivers Openness And Accelerator Portability At ModCon 2026

Forbes Published Aug 20, 2026 Reviewed Aug 20, 2026 ✓ Reviewed by citations.press editors
Modular Delivers Openness And Accelerator Portability At ModCon 2026
Qualcomm acquired Modular for about $3.1 billion in stock, representing 18 million shares.
about 3.1 $ · transaction value Qualcomm’s quarterly SEC filing, SEC filing
Chris Lattner, former Modular CEO, became Qualcomm’s executive vice president for advanced AI software and platforms.
1 · executive vice president Modular CEO Chris Lattner, executive vice president
AMD’s MI355X chips delivered roughly 90 % faster throughput than Nvidia’s B200 at about half the hourly cost, resulting in more than 45 % lower total cost of ownership.
roughly 90 % · throughputabout 50 % · hourly costmore than 45 % · total cost of ownership Modular, presented performance
Using Modular, five engineers brought up Amazon’s Trainium chip in under five months, totaling roughly 25 engineering‑months versus several thousand traditionally.
5 engineers · engineers involvedroughly 25 engineering-months · engineering effort Modular, presented
Modular created the Google TPU version largely on its own in about four months.
about 4 months · development time Modular, presented
Qualcomm’s AI200 went from hardware access to running a model in two weeks.
2 weeks · time to run model Modular, presented
MiniMax runs its flagship model on a dedicated Modular deployment serving billions of tokens per minute.
billions 1 tokens per minute · tokens per minute Modular, reported
Hippocratic AI reports better than 30 % gains for its voice agents run via Modular.
more than 30 % · performance gains Modular, reported
Modular Cloud has been serving live traffic for months on OpenRouter under the stealth name ModelRun.
Modular, disclosed

Following Modular’s $3.1 billion purchase by Qualcomm, the ModCon 2026 conference presented several important updates. The company fully open-sourced its Mojo programming language and compiler, addressing concerns about platform neutrality. Modular Cloud, an AI inference service, is now generally available, supporting chips from suppliers including Nvidia, AMD, Google, AWS and Qualcomm. Modular presented examples of the platform drastically reducing cost and time-to-market for AI chips. Modular Cloud also promises substantial cost savings for AI workloads and positions Qualcomm at the arbitration layer of AI computing. While some reservations about the MAX license and Modular Cloud's proprietary nature remain, ModCon offered a credible alternative to existing AI software ecosystems.

I spent Tuesday in San Francisco at ModCon 2026, Modular’s developer conference and its first event since Qualcomm acquired it just a few weeks ago. Going in, I posted on X that when I ask people about Modular, I get “bookend” responses with no middle ground: One end believes Modular is building the most important software layer in AI, while the other points to the graveyard of companies that tried to break the AI software moat. I left the conference believing that the skeptics now have the harder case to make.

Qualcomm gets plenty in return for the $3 billion-plus in stock it spent to buy Modular: day-one software for its own chips from the datacenter out to the edge, a new services business and a position in the layer that decides where AI workloads run — even when a Qualcomm rival’s chip wins. (Disclosure: Qualcomm is a client of my firm, Moor Insights & Strategy, as are other companies cited here, including AMD, AWS, Google, Microsoft and Nvidia. I attended ModCon at the company’s invitation. The analysis is entirely my own.)

ModCon delivered three things the industry has been waiting on. First, Qualcomm legitimated its open source commitment within three weeks of closing the acquisition, and with an actual license rather than a promise. Second, Modular showed working evidence that its platform can radically lower the cost of bringing new AI chips to market. Third, it lowered the cost of compute for the token consumers who can arbitrate not only across models, but now chips. Real unknowns remain, and I’ll name them, but this was a stronger day than I expected — and I already expected a good one.

It’s worth recalling the short timeline of this acquisition. Qualcomm announced the Modular deal on June 24, an all-stock transaction announced at roughly $3.9 billion, and closed it on July 28. Qualcomm’s quarterly SEC filing puts the actual transaction value at about $3.1 billion for 18 million shares, since all-stock deal math moves with the stock price. Modular CEO Chris Lattner, co-creator of the LLVM compiler infrastructure and Apple’s Swift language, is now a Qualcomm executive vice president running advanced AI software and platforms.

Every acquisition of a neutral platform by a hardware vendor raises the same question: Does it stay neutral? My colleague Matt Kimball framed the stakes in his article from late June about Qualcomm’s investor day: If Modular’s openness narrowed toward Qualcomm silicon, it would become “a lock-in play wearing open-source clothing.” When my colleague Bill Curtis wrote about the deal 10 days later, he set two tests: “when a [Qualcomm] Hexagon [NPU] backend ships, and whether MAX stays open to competing silicon.” I asked the same question in my Week Ahead video going into the event, and the developer community was asking it right up to the week before.

ModCon answered on both fronts. The Hexagon question was resolved in silicon: Qualcomm’s new datacenter chips are built on its Hexagon processor lineage, and they now run the Modular stack, with a Google Gemma model working on the AI200 within two weeks of hardware access. The openness question was resolved through licensing. Releasing under an Apache 2.0 license isn’t something that can be quietly walked back, and Modular’s announcement commits in writing to optimizing for hardware “including hardware that competes directly with Qualcomm Technologies’ platforms.”

Qualcomm’s Amon left no ambiguity with what he said on stage, and when I sat down with him and Lattner afterward, he repeated it to me directly: “We did not acquire Modular to make it a Qualcomm-only software.” He framed the moment to me as Android and Kubernetes arriving for AI at the same time. Lattner told me the pitch that won him over was achieving Modular’s mission faster on a bigger, more open platform. I believe that there were multiple bidders with higher offers, making openness the thesis of the deal, not a concession. Trust was what Modular actually needed to ship, and in the software world a license is the only durable form of it. You can’t un-open-source a compiler.

The obvious question is why Qualcomm would pay billions in stock for a software company and then give away its most valuable asset three weeks later. The direct answer: Qualcomm’s datacenter chips now have production software from day one, and the software question that had dogged every Qualcomm datacenter conversation is now answered. It also gets a software layer it can take across its entire edge AI portfolio. As Matt Kimball argued in June, a credible software layer de-risks Qualcomm’s entire hardware roadmap in one move. Modular Cloud gives Qualcomm usage-based services revenue it never had. And in the bear case, Qualcomm still walks away with one of the strongest AI software teams ever assembled.

The indirect benefits could be bigger. Modular Cloud puts Qualcomm at the arbitration layer of AI computing: It earns revenue on every token served and sees demand across every vendor’s chips, including in deployments Qualcomm hardware hasn’t won. Opening the stack also levels the playing field exactly where Nvidia has an advantage in software. This shifts competition to performance per dollar per watt — the fight that Qualcomm spent two decades training for in mobile. And it completes the company’s developer-first repositioning that Amon walked me through as he discussed its acquisitions from Edge Impulse to Arduino to Modular. One stack for everything from earbuds to datacenter racks makes every Qualcomm socket more valuable, and the edge is where Qualcomm is strongest.

A hat tip to Cristiano Amon is in order. He has spent 30 years at Qualcomm making big bets that the market doubted, and I remember when few people believed in the diversification strategy that now has the company tracking, by its own projection, to be the largest automotive chip supplier within two years. Paying more than $3 billion in stock for a crown jewel and then fully open-sourcing a layer in week three takes a strategic self-confidence that most acquirers never muster.

I believe the ModCon 2026 announcement that will matter most commercially is Modular Cloud. Think of it as air traffic control for AI: the cluster, not the chip, is the computer, and the platform decides which silicon serves each request. Developers see one standard interface; Modular handles the hard engineering underneath. Routing requests to the right model is a crowded field, one that Nvidia entered days earlier with its NeMo Switchyard router. Routing them to the right hardware at production quality is the part almost nobody else has attempted. I posted the full value-proposition slide live while the keynote was underway, and the more I reflect on it, the more I believe “audacious” is the right word for it.

What separated this launch from a typical debut is that it wasn’t really a launch, because the product is already in service at scale. Modular disclosed that it has been serving live traffic for months on OpenRouter, a popular marketplace for AI models, under the stealth name ModelRun; Modular says its endpoints consistently ranked at or near the top for speed, a story that independent benchmarking firm Artificial Analysis echoes. AI model maker MiniMax runs its flagship model on a dedicated Modular deployment serving billions of tokens per minute; healthcare AI company Hippocratic AI reports better than 30% gains for its voice agents run via Modular; and Jane Street, a trading firm for whom microseconds are money, is also a paying customer. This isn’t a research project.

The real kicker for the Qualcomm/Modular platform approach is the radically lower cost of bringing up new chips. Enabling a full AI software stack on new silicon traditionally takes hundreds of engineers and engineer-years; that is the source of CUDA’s gravity in the market. Modular’s counterargument is arithmetic. Using Modular, five engineers brought up Amazon’s Trainium chip in under five months, roughly 25 engineering-months against the several thousand a traditional effort consumes. Modular created the Google TPU version largely on its own in about four months. Qualcomm’s AI200 went from hardware access to running a model in two weeks. These are vendor numbers, so I treat them with appropriate skepticism. But even heavily discounted, a tenfold reduction in enablement cost changes who can afford to ship credible AI chips. At ModCon, AMD committed on stage to a multi-generation partnership, and chip startup d-Matrix signed on, too.

During the event, Modular put hard numbers behind Modular Cloud’s performance. Holding deployment size and response-time targets constant, the company showed the same model on AMD’s MI355X chips delivering roughly 90% faster thoughput of an Nvidia B200 setup — at about half the hourly cost. That’s more than 45% lower total cost of ownership. I called the numbers “stunning” live. Mind you, the test details weren’t published, and none of it has been validated by Signal65, the independent benchmarking lab I co-founded. So please treat every figure here as a credible vendor claim awaiting third-party verification.

This engineering direction aligns with everything I’m hearing about memory use. I posted during the event that, in the context of the ongoing memory crunch across the industry, at least a dozen chip design companies have told me they’re rearchitecting to use less memory, and that they believe memory supply will stay short for years. Qualcomm’s datacenter chips aim squarely at that constraint: The AI200 is slated to ship this year with less expensive memory per card than typical GPUs using HBM, and the AI250 arrives in 2027 with a memory-first design that, as Yahoo Finance" href="https://moorinsightsstrategy.com/research-notes/patrick-moorhead-discusses-qualcomm-and-micron-on-yahoo-finance-june-25-2026-broadcast-analysis/" rel="nofollow noopener" target="_blank">I discussed on Yahoo Finance, improves bandwidth and energy use without needing the advanced chip-packaging capacity that’s in shortest supply. A software layer that routes AI work across many kinds of chips makes that hardware diversity deployable rather than merely announced.

Here’s a second-order effect nobody else is talking about. AI datacenter builders are financing tens of billions of dollars of infrastructure, and lenders price that debt on risk, which today means anything other than the consensus Nvidia GPU. A credible cross-chip software layer changes that math. If the same models and applications run smoothly across Nvidia, AMD, and purpose-built accelerators, then a mixed fleet becomes a lower-risk asset, and lower risk should eventually mean cheaper capital for the people building AI factories. I believe the abstraction layer is quietly a financing story, too, not just an engineering one.

I’d be doing readers a disservice if this article read like a final victory lap, so let’s be candid about the areas for further consideration, starting with the license rather than the press release. Mojo is now open source; but MAX, the layer above it, is source-available under a community license, and those are not the same thing. The old device limits are gone, which is real progress, but conditions remain on telemetry, branding and building substitutes, and source access is not the same as community governance.

The bigger customer risk could be a lock-in moving up the stack: Modular Cloud, the layer that decides which chips get your workloads, is proprietary. Apache 2.0 constrains the programming language; it doesn’t constrain the cloud. The alliance program’s actual terms, promised by year end, will show whether partners get real governance or merely a logo slide.

We saw a lot of work completed, and it’s clear there’s more to come. And with that work, execution needs to be flawless. Per Modular’s own materials, Qualcomm’s Cloud AI100 is serving through Modular Cloud today, while the Google and Amazon chips run models but aren’t in the cloud service yet. Production is expected within months, but it’s worth remembering that announced support, production readiness and cloud availability are three different milestones. I believe roughly 100 engineers support all of it, which is either the best leverage in infrastructure software or an area of fragility waiting for a bad quarter. If they can pull off the rest of it with 100 engineers, all of them should be given raises.

Meanwhile, Nvidia isn’t standing still, and its advantage is rooted in a decade of tooling and industrywide habits. Modular’s cost claims in comparison to CUDA need independent validation, and smart buyers will pilot real workloads on at least two kinds of hardware and measure cost, speed and engineering effort before standardizing on anything. They could also maintain an exit path by keeping model weights, prompts and deployment automation portable to ensure that leaving remains a cheap alternative.

Four and a half years ago, Modular bet that AI inference would become the dominant workload, that AI chips would diversify and that the industry needed one software layer to make that diversity usable. Every one of those bets now looks like consensus, and at ModCon the company showed a production platform, a public cloud, an open compiler and six chip vendors to back it up.

For Qualcomm, the multiple types of payoffs from this acquisition is what makes it a smart move by Cristiano Amon. The floor is adding a world-class software team and day-one software for its own datacenter chips, and moving out to the edge. The ceiling would be owning the neutral layer the AI industry runs on, earning money on every token served no matter whose silicon wins the socket. To enable this, Amon traded a lock-in nobody would have trusted for a position everyone can use, and he accepted the strange math at the deal’s core: Qualcomm now has to help its competitors succeed to get paid. Most CEOs can’t stomach that trade; the smart ones recognize it’s the only one that works if you want to be a trusted platform. I confess that I’m rooting for Qualcomm and Modular to pull it off, because while I have nothing but respect for what Nvidia has accomplished, the industry needs more platform diversity.

But platform trust is never won permanently, and the potential pitfalls I’m watching are the MAX license, the alliance terms, independent validation of Modular’s numbers and whether the bring-up economics hold for the next chip. But the burden of proof has shifted from the people who believe in an AI software alternative to the people who don’t.

This article was originally published by Forbes ↗. citations.press indexes the source-backed facts above and links to the original. Something wrong? Corrections policy · Report an error