TL;DR: The training era was centralized: a few hyperscalers, giant clusters, one bottleneck called GPUs. The inference era is distributed: intelligence running where the data lives, on systems that cannot crash, delivered over infrastructure nobody trusts by default. Six companies I follow, NVIDIA, Nebius, Cisco, NXP, BlackBerry, Akamai, and Palantir, are not six separate stories. They are six layers of one stack.

Two eras

Every technology has a training phase and a deployment phase. The training phase centralizes: talent, capital, and compute gather in a few places, and one supplier captures most of the value. For AI, that was NVIDIA and the hyperscalers. The whole market watched one bottleneck.

The deployment phase distributes. Intelligence has to run where the data is generated: in the car, on the factory floor, in the hospital, at the retail edge. It has to answer in milliseconds, it cannot crash, and it has to do all of that inside systems nobody fully trusts. That is a completely different engineering problem from training a model, and it needs a completely different stack.

This is the pattern I keep coming back to. I call the method "what's missing in the stack": start from a working thesis, look at the adjacent layers, and ask which piece the map doesn't have yet. The inference stack is that map, filled in.

The layers, bottom to top

No layer has a single winner. Every layer has several credible players, and we actively follow only a few of them. In the diagram above, the bright names carry DalalBytes research; the dimmed names are on the watchlist but uncovered. Intel sits in the edge-silicon layer, and it belongs there. And each layer exists twice: once in the cloud, once at the edge. The diagram shows the fork.

OPERATIONAL INTELLIGENCE · SPANS CLOUD + EDGE Palantir Databricks CLOUD EDGE DELIVERY · API FRONT DOOR Akamai Cloudflare DELIVERY · ON-DEVICE + FAR EDGE on-device inference PLATFORM · CONTAINERS Kubernetes PLATFORM · REAL-TIME OS BlackBerry QNX VxWorks NETWORK · DATACENTER FABRIC Cisco Arista Broadcom NETWORK · 5G / TSN / INDUSTRIAL Private 5G · TSN SILICON · CLOUD INFERENCE NVIDIA Nebius AMD SILICON · EDGE INFERENCE NXP Intel Qualcomm SECURITY HYPERSCALERS full-stack alternative AWS Azure GCP rent GPUs, not buyers own fabric (white-box) own CDN front door own AI platforms own edge extensions compete at every layer sovereign stacks exist for customers who want out of this column DalalBytes covered Tracked, not covered One stack, two deployments. Intelligence spans both.
The inference stack, forked: every layer has a cloud and an edge instantiation, operational intelligence spans both, and hyperscalers compete at every layer as the full-stack alternative.

Compute: NVIDIA and Nebius. Somebody has to sell the raw intelligence. NVIDIA sells the GPUs, the networking, and increasingly the full AI factory blueprint. Nebius sells the same capacity as a neocloud, for everyone who is not a hyperscaler. AMD is the credible second source, and the hyperscalers themselves are quietly becoming compute vendors to everyone else. This is the layer the market already prices as AI.

Fabric: Cisco. Compute is useless until it is stitched together. Cisco's Hyperfabric and Silicon One are the connective tissue of the AI factory: the switches, the silicon, and the management plane that turn a pile of GPUs into a working cluster. Arista owns the hyperscale DIY end of this market, Broadcom the merchant-silicon end. The September collaboration with Palantir and NVIDIA extends Cisco's fabric into sovereign deployments. Boring, essential, and very hard to rip out once installed.

Edge silicon: NXP. Inference at the edge runs on microcontrollers and processors in cars, factories, and industrial equipment. NXP owns an enormous share of that footprint. Intel is in this picture too, through its edge portfolio, AI PCs, and foundry ambitions; Qualcomm through mobile and automotive. As those chips absorb more AI workload, the silicon content per device grows. The market still prices NXP as an auto-cycle stock. The inference era may reprice it as edge-compute real estate.

The deterministic OS: BlackBerry QNX. This is the layer most investors skip, and it might be the most defensible. When software steers a car or runs a surgical robot, it cannot crash, freeze, or get hacked. QNX is the real-time operating system certified for exactly that, and it earns royalties per unit shipped. Wind River's VxWorks plays in the same arenas. Every layer above it can be swapped. The certified OS at the bottom almost never is.

Distributed delivery: Akamai. The first cloud player is now the inference cloud: thousands of edge locations sitting closer to users and devices than any hyperscaler region. Cloudflare is building the same geography from the developer side up; Fastly holds pieces of it. When agentic AI needs low-latency answers at planetary scale, the request does not travel to a data center. The data center travels to the request. Akamai already built that geography for the web. Now it serves it to AI.

Operational intelligence: Palantir. At the top sits the software that turns operations into decisions. The ontology models a company's real world, suppliers, machines, patients, vehicles, and the AI reasons over it. Databricks is coming at the same problem from the data side. The September collaboration puts Palantir's ontology on Cisco's factory infrastructure with NVIDIA's models underneath. Hardware, fabric, and software in one validated stack.

Security is not a layer

Notice what is missing from the layer cake: security. That is deliberate. In the inference era, security is not a product you buy at the end. It is the thread through every layer: sovereign control over data and models, safety certification in the OS, zero-trust in the fabric, attested delivery at the edge.

This is why Jensen Huang keeps naming cybersecurity as AI's next major use case. Every enterprise deploying agents is really asking one question: how do I let the machine act without losing control? The companies that answer it at each layer collect the toll.

Cloud and edge: the stack forks

So far this reads as one vertical stack. Technically it is two. Every layer below intelligence has a cloud instantiation and an edge instantiation. Cloud inference runs on GPU clusters behind an API, stitched with datacenter fabric, orchestrated on Kubernetes. Edge inference runs on NPUs and microcontrollers, over 5G and time-sensitive networking, on a real-time OS that cannot crash.

The layers do not change. The deployment does. That is why operational intelligence sits above the fork: the ontology is the one layer that spans both, letting the same model of the world reason over cloud data and edge data together. Palantir's architecture is literally this picture, cloud and edge feeding one ontology.

And the hyperscalers do not sit in any layer. AWS, Azure, and GCP each run the whole stack themselves: rented GPUs, white-box fabric, their own CDN front door, their own AI platforms, their own edge extensions. They are the full-stack alternative, competing at every layer simultaneously. The sovereign AI announcements only make sense against that backdrop: Cisco, Palantir, and NVIDIA are assembling the enterprise answer for customers who want out of the hyperscaler column. Akamai's entire pitch is being the neutral edge, the one that is not a hyperscaler.

For the investor, the fork doubles the hunting ground. Every toll booth exists twice, and the edge-side booths are the ones the market has barely started pricing.

Case studies: the stack in action

Frameworks are cheap. Here are three places the stack is already working, each one traceable to a real announcement.

1. The sovereign AI factory (Cisco + Palantir + NVIDIA, September 2026)

Cisco, Palantir, and NVIDIA announced a collaboration for sovereign AI infrastructure: Palantir's Ontology for Cybersecurity delivered through Cisco's Secure AI Factory with NVIDIA, published as the Palantir Sovereign AI Operating System. A validated reference architecture combining compute, networking, storage, security, and observability for governments and regulated enterprises.

Read it as a stack trace. NVIDIA supplies the models (fine-tuned Nemotron) and the factory blueprint. Cisco supplies the fabric, the security, and the observability plane. Palantir supplies the ontology that lets a government or a bank reason over its own operations without surrendering control of the data. Three layers, one validated product, announced by Cisco's product chief.

The investment read: enterprises do not buy AI components, they buy validated stacks. Whoever owns the validated reference architecture collects the toll. Watch for who gets named in the next three such announcements.

2. The software-defined vehicle (Uber + QNX, plus NXP silicon)

On its Q2 FY27 earnings call, BlackBerry confirmed that Uber selected QNX as the software foundation for its next generation of vehicles. Separately, the NVIDIA engagement around QNX continues on its own track. The car is becoming a rolling inference endpoint: dozens of NXP processors running a deterministic QNX core, making millisecond decisions that cannot fail.

Read it as a stack trace. NXP's silicon content per vehicle keeps rising as cars absorb driver-assistance and autonomy workloads. QNX earns a royalty on every unit that ships with it inside. Over-the-air updates ride on distributed delivery infrastructure. Fleet-scale learning closes the loop back into operational intelligence. One vehicle, four layers of the stack, and two of them are annuities paid per car.

The investment read: per-unit royalties plus rising silicon content turn the car into a compounding annuity. The market still values these as auto-cycle plays. The stack says they are inference endpoints.

3. NVIDIA's own supply chain (Palantir SAIOS on Dell and Cisco infrastructure)

NVIDIA and Palantir extended their sovereign AI operating system reference architecture to supply chains: Nemotron models post-trained on proprietary operational data, optimized with NVIDIA's own software, deployed on infrastructure from Dell and Cisco, with hosted options including Nebius. The first deployment runs on premises at NVIDIA itself, on its most sensitive operational data.

Read it as a stack trace. When the company that sells the picks and shovels runs its own mine on your software, the reference architecture stops being a partnership slide and starts being the enterprise default. Compute, fabric, and intelligence in one validated loop, with the infrastructure vendors collecting their share underneath.

The investment read: reference architectures compound. Each validated deployment lowers the sales cost of the next one. The winners are the names that appear in the architecture diagram, not the press release.

The pattern only appears from above

Here is the part no machine did for me. Each of these companies has a fine standalone research report on this site. None of those reports mentions the pattern, because the pattern does not live inside any single company. It appears only when you hold all six in your head at once and ask what they are together.

That is still the human job. The machine researched every layer in hours. Seeing the stack took the years of building the map: the edge-AI hunt that started in 2023, the Palantir cycle that taught the ontology, the QNX royalty insight that no screen would surface. AI compressed the research. It did not compress the seasoning.

Where the money hides

The market is efficient at pricing the obvious layer. Everyone knows NVIDIA is AI. The inefficiency lives in the layers the market has not repriced yet: an operating system royalty stream still valued like a legacy software business, an auto chipmaker sitting on edge-compute real estate, a content delivery network that quietly became inference infrastructure.

That is the hunting ground. Not the layer everyone is staring at, but the adjacent ones the stack cannot work without. Find the toll booths the market hasn't noticed, check them against the pillars, and let the slow filter do its work.

The training era made one company the trade. The inference era makes the stack the trade.

Research and opinion, not investment advice. Do your own due diligence before investing.