The Rise of Compute Capital Markets | Pantera

The Rise of Compute Capital Markets

September 9, 2026 | Jay Yu

Source: Sherwood [11]

 

Compute and datacenter infrastructure spend has become a trillion dollar category, rivalling US consumption as a major pillar of US economic growth. Yet, for all this demand for compute as AI agents trickle into the real world economy, the process of actually acquiring GPUs is surprisingly old-school. While marketplaces like SF Compute, Vast AI, and Runpod all exist, much of the actual buying and selling of GPU nodes is actually run through a collection of group chats, over-the-counter (OTC) brokers, and bespoke bilateral agreements.

 

 

Today, GPU compute is still early in its financialization lifecycle [1]. Like electricity, compute is not a totally fungible asset – it is bound by the constraints of SKU, time, and location. But fast forward 5-10 years, compute may well be come a commodity asset just like power or oil – a critical global resource underpinning economic output in the AI era. In this essay, we will investigate the process of how this occurs – what are the lessons from electricity that we can take to understand compute markets structure, the landscape of products being built in the space, as well as both the opportunities and open questions of commoditizing compute.

 

1 – Learning from Power Markets

1.1 – The Structure of Power Markets

To understand how the structure of compute markets may evolve, we should turn to the markets of their underlying electrons. Just like compute, electricity is a heterogeneous, temporal, and hub-based asset that underpins global economic growth. Beginning in the 1990s, US electricity systems gradually became privatized, and vertically integrated utilities such as generation, transmission, and distribution becoming increasingly unbundled [2]. Independent grid operators began servicing various parts of the country, and the largest interconnects, such as PJM-West (East Coast), ERCOT North (Texas), CAISO SP15 (SoCal) eventually became tradable commodities on ICE and CME and a collection of benchmarks for the asset class [3].

 

At the physical layer, electricity markets essentially have a “Grid-Operator-Node” structure. At the top are a collection of physically separate regional grids. For each of these regional grids, there is a collection of system operators – Regional Transmission Operators (RTOs) and Independent System Operators (ISO). These operators run a series of local substation nodes. A city such as San Francisco, for example, could be serviced by a number of these substation “nodes” (eg. Embarcadero, Larkin, Mission). The pricing at each individual node varied according to an algorithm called “Locational Marginal Pricing” (LMP) which ran a constrained optimization over generation dispatch and customer demand, physical transmission costs, and power flow physics [4].

 

Image Source [5]

 

The LMP algorithm became the key link between the physical delivery of a Watt of power with the tradable electricity benchmark such as PJM-West. The “hub” and index prices for power are essentially weighted averages over baskets of the end LMP nodes. The largest of these eventually became stable and liquid enough to be tradable on exchange venues like CME and ICE, and became the reference prices for the entire asset class.

 

1.2 – Hypotheses on Compute Markets

Looking at this evolution of power markets in the last 30 years, there may be several hypotheses about compute markets that we can draw.

 

First, compute markets may follow a similar structure to these power markets. On the physical settlement side, perhaps the best analog to the “Grid-Operator-Node” structure of power markets is “Hardware-Vendor-Cluster.” Just as power markets have physically separate grids, compute markets may be defined by the physical hardware class. H100s, H200s, B200s, B300s and other chip classes have their separate (although interrelated) benchmarks. At the “operator” level, we have a variety of different compute vendors, such as AWS, Nebius, Coreweave, SF Compute, Ornn, each offering different pricing mechanisms for a particular clusters differentiated by time/location/SKU similar to power market nodes. As for LMP, the closest analogy may be a scheduling/routing algorithm that emits dynamic instance prices based on a particular SKU availability.

 

The “benchmarks” that won in power markets and became the reference index for the asset class (ie. PJM-West, ERCOT-North, CAISO SP15) were largely the ones that had the most liquid physical delivery infrastructure – similar to the structure of oil and other commodity markets. As we see numerous benchmarks emerge in the compute space, perhaps the ones that win will be those tied to the most liquid physical settlement venues.

 

Moreover, in power markets, although principals (generators on shortside, load-consuming entities on the longside) sometimes directly hedged on the power exchange, most of the “principal” activity goes through OTC desks run by banks and trading firms that swap nodal settlement for hub settlement at a spread and hedge their exposure. This market dynamic may also play out in compute markets, where the actual “principals” of compute (neoclouds and AI companies) may prefer to go through brokers for their SKUs rather than directly hedge on these markets.

 

Finally, compute markets may see much more basis risk than power markets. While wholesale power rates are mandated by the FERC to be transparent under the Federal Power Act, there is no such transparent mandate for compute markets [6]. The result here is that you can only create dynamic prices (and indices) based on your own orderbook and the orderbooks of neoclouds that you partner with via revenue sharing and data acquisition deals.

 

2 – The Structure of Compute Markets

2.1 – The Inference Demand Stack

Just as power markets exist because of the industrial applications of electricity, compute markets exist because of the demand for tokens, and in particular for inference tokens.

 

Source: Original Content

 

Today, the inference stack can be largely be seen as a having three distinct layers that form the structural longs and shorts for compute:

 

  1. Neocloud layer: players such as Nebius, Coreweave, and others operate physical datacenters and form the sell-side for GPUs.
  2.  
  3. On-tap layer: developer platforms such as Fireworks and Baseten then take these baremetal environments into “rich” GPU environments that a developer can run tasks or get inference tokens on-tap. This layer is a buy-side for GPUs.
  4.  
  5. Application layer: applications such as Cursor, Perplexity, Rime etc. make use of these inference platforms to deliver an end-product to users and businesses. This layer is a buy-side for tokens (which in-turn requires GPUs)

 

While each player in this ecosystem has their own slightly different strategy of managing GPUs, largely speaking the neoclouds are the structural short-side for the GPU markets, whereas the on-tap and application layers are the long-side. On the other hand, a hyperscalar such as Amazon or Google operates different inhouse products at every layer of this stack. As for the “margin flow,” a rule of thumb is that when one of the top-layer apps spends $100 in tokens, ~$45 goes to the on-tap layer, and $50 to the neocloud/GPU layer. The remaining $5 is on a routing layer such as OpenRouter.

 

2.2 – Compute Capital Markets Structure

 

Source: Original Content

 

These underlying structural “tokenomics” of the AI principals affect the shape of how compute capital markets get formed. One way that the structure of compute markets may evolve, is that “AI principals” on both the long (dev platforms and applayer) and short (neoclouds) side both interface with compute brokerages and OTC desks such as SF Compute, Runpod, Compute Exchange for particular SKUs. Then, the compute brokerages and OTC desks will manage inventory and absorb basis risk between the SKU demanded by the end token consumer and the generic H200 by hedging on a compute exchange venue such as Architect or Pluto. These exchange venues’ prices are in turn based on a benchmark built based on weighted averages from neocloud and OTC platforms’ orderbooks that these platforms partner with.

 

2.3 – NVIDIA is the Central Bank for Compute

In addition to this, NVIDIA has also recently announced that it will allow its AI factory compute to become an “investable asset class” [7]. Within this compute financialization stack, one way to view NVIDIA’s involvement is as a “central bank” given its monopoly on the GPU tech stack. A central bank typically has several key goals: (1) manage inflation in an economy, (2) allow for full employment, (3) be a lender of last resort. NVIDIA arguably is playing a role towards all of these goals:

 

  1. Managing Inflation: the analog of “inflation” in the GPU economy is arguably the depreciation cycle, especially relative to SOTA chips. By being able to control product releases (eg. Vera Rubin), NVIDIA is able to have some degree of influence over how quickly or slowly GPUs depreciate over time.
  2.  
  3. Full Employment: NVIDIA wants “full employment” of its GPUs – if utilization of GPUs are very high (due to high demand for tokens), this means that people will want to buy more units of its GPUs. Therefore NVIDIA is also incentivized to increase the financialization and make GPU-hours more liquid, such that its utilization can increase.
  4.  
  5. Be a lender of last resort: NVIDIA has also announced that it will provide up to 25% residual value support, which can be seen as a form of bailout/insurance guarantee of GPU value in the case that neoclouds and other GPU providers run into liquidity crunches.

 

3 – Compute Market Opportunities and Challenges

3.1 – The Product Shapes of the Compute Stack

As the financialization layer of compute markets continues to grow, there seems to be several main products that can be built in the space, each with their advantages and challenges.

 

Product 1 – Physical Settlement. This refers to the process of actually delivering physical GPU units to the end consumers that actually need them, and includes players such as SF Compute, Hyperbolic, Vast, Runpod, Compute Exchange. This is the product layer closest to the metal and actually interfacing with the AI principals mentioned above. Because of the heterogeneity in compute SKUs, many players here start out simply as brokers, taking commission on matching GPU demand with individual neoclouds [8]. Eventually, the end goal here is some form of a “spot exchange” for compute, but consistently controlling for quality at the physical delivery of the GPU metal is actually a hard problem, especially for players whose supply comes from a decentralized network. It is possible that for this layer to mature, we need to see a Moody’s like ratings agency that guarantees the quality of the underlying deliverable.

 

Moreover, despite this layer being hypercompetitive, physical settlement is arguably the layer with the most durable moat in the long term. Looking back at cases such as oil and power, we see that it is typically the places with the most liquid physical settlement venue that the financialization layers (eg. index, exchange, and lending) typically develops around [9]. So the prize of getting physical settlement right is huge.

 

Product 2 – Index. This layer refers to the “index curves” that are being built on top of compute, such as those by Ornn, Silicon Data, Compute Desk, Semianalysis and others. The construction of an index is a vital step in allowing an asset class like compute to be further financialized and tradable. Arguably much of the current “compute markets” narrative was kickstarted as compute indices as a product layer began to mature, and many major exchanges such as CME and ICE announcing partnerships with these indices.

 

Image Source [10]

 

However, today there still is an extremely large spread between many of the compute indices and the individual GPU marketplaces. There may be two key factors to this price differential. Firstly, there is dispersion in the underlying product, as there is a huge difference in product reliability, interruptibility, and contract terms across each platform. Secondly, because the orderbooks underneath all of these different compute indices are different, and there is no transparency mandate here in the same way that there is for other asset classes (eg. FERC mandates for power). Many of these compute indices today are simply buying orderbook data in bulk from neoclouds or doing partnerships/revshares to bootstrap the underlying datasets. Moreover, the index layer in and of itself is hard to monetize compared to the exchange and physical settlement layers. Thus, despite being able to make for splashy announcements and being an important layer in the compute markets stack, the index layer alone faces a double-squeeze from both the physical layer (from which it sources prices from) and from the exchange layer (which demands a revenue cut).

 

Product 3 – Derivative Exchange Venues. The third layer in the stack is the exchange venue layer (typically cash settled futures), which provides a hedging venue for compute, based on an index product, such as those being developed by Liquid Compute and Architect. While this may be the product that is arguably the most exciting and easily monetizable in the long run, it is still in its early stages, without large amounts of trading volumes. Moreover, because the tradable product on cash-settled venues is a generic H200 instance rather than a particular SKU on which you can run inference, most of the players interfacing and trading on these exchanges may just be OTC desks and compute brokers that wish to hedge the inventory that they hold on their balance sheet, rather than the end AI principals.

 

Product 4 – Financialization Vehicles. A looser fourth category of products that can be built in the compute space is vehicles for this financing to take place. For example, there may be lending protocols, vaults, and synthetic stablecoins such as USD.AI that help to finance the physical buildout of datacenters and neoclouds. We may also see risk transfer and insurance layers appear too, as a way to smoothen out the basis risk in the GPU economy.

 

Looking at all of these above layers, we can see that as a category, the end shape of compute as a category is fairly clear – the winner is a fullstack player that is able to allow for robust physical settlement of GPUs, provide an index as an emergent property of its physical orderbook, and uses this index to build up a futures exchange where compute can be hedged as an asset class. But where the disagreements happen is over the sequencing of the products in the space – will the winner start with physical settlement first, or go with an index, or start with a trading venue first?

 

3.2 – The Frontier of Compute Tokenomics

Thus far, we’ve primarily talked about “compute markets” purely at the GPU level. But we can also zoom out a little and look at compute markets amongst broader “tokenomics” of compute – from a historical perspective, from a value capture perspective, and from a game-theory perspective with the frontier labs.

 

Historically speaking, “compute marketplaces” are not necessarily a totally new idea. In 2023-2024, a wave of players such as SF Compute and Hyperbolic were already describing this idea, alongside decentralized players such as Akash and IONet. What’s interesting is that many of the players didn’t end up just staying at the GPU marketplace layer. Players such as Hyperbolic also became inference providers to directly sell tokens, while SF Compute began to contract and operate clusters in a first-party way.

 

Indeed, from a value stack perspective, value flows from AI applications (eg. Cursor, Harvey, Granola) to routing layers such as OpenRouter, to inference providers (such as Fireworks and Baseten) to neoclouds and marketplaces (Nebius, Coreweave, SF Compute) and down to datacenter operators. Staying right in the middle of the stack may turn out to be a squeeze from both ends – you have to move either up or down to maintain more robust margins.

 

Finally, there is an extremely interesting game theory perspective of token markets with the three “elephants in the room” – (1) NVIDIA on the one hand controlling supply depreciation, (2) frontier labs (eg. Anthropic and OpenAI) releasing frontier models and spiking up token demand, and (3) open source models (eg. Kimi, GLM, DeepSeek, Qwen) placing downwards pressure on inference margins. The combination of these 3 elephants gives structural catalysts on both the long and short side of GPU and token markets, and any of these three elephants have the power to release major products that rewrite the entire pricing curve. This is because a discretionary event (eg. the release of GPT 5.6) can have cascading ripple effects along the entire value stack of AI tokens – causing fluctuations in token prices and GPU metal prices. Because of this, arguably a company that operates both token inference as a business and a GPU metal marketplace under the same roof is extremely economically rational, as it offers a hedge against the change in margin profile of these two layers in the stack.

 

Conclusion

Looking back at history, many major commodity markets – such as oil and power – started off roughly where compute markets sits today: with bilateral, opaque trades run by middlemen. But gradually, physical settlement evolved into a set of protocols and financial primitives, such as indices, transaction standardization, and quality auditing, before turning into a fully financialized asset class.

 

Today, we’re seeing this transformation live in compute. Already, we see a diverse set of projects experimenting with financialization primitives such as indices, cash-settled exchanges, and lending and insurance products. At the same time, we see the physical GPU marketplace and cloud providers moving to capture the entire value stack of tokens. As the fundamental asset underpinning the entire AI economy, compute is increasingly evolving from a piece of infrastructure into a financial asset in its own right. We will likely see multiple large unicorn winners in the space, each spanning a different layer of the market and taking form in a different shape – from physical provisioning to brokerages to lending and risk management, and with composability with both DeFi and TradFi capital markets. Compute may be the first major new physical commodity to emerge in decades – and we may be just in time to witness its transformation from a group chat to a full asset class.

 


IMPORTANT DISCLOSURES

This document is made available by Pantera Capital Partners LP (“Pantera”) for informational and educational purposes only. It does not contain all information pertinent to an investment decision. Nothing in this document constitutes an investment recommendation or an offer of investment advisory services. This document cannot be relied upon in making an investment decision. Nothing contained herein constitutes an offer to sell, or a solicitation to buy, any securities. This document contains information believed to be reliable, and has been obtained from sources believed to be reliable, but no representation or warranty is made (express or implied) of any nature, nor is any responsibility or liability of any kind accepted, with respect to the fairness, accuracy, completeness, or reasonableness of the information or opinions contained herein. Forward-looking statements should not be relied upon. There is no guarantee that investments in any instrument described herein will be profitable – all investments carry the inherent risk of total loss. Analyses and opinions contained herein (including market commentary, statements or forecasts) reflect the judgment of the author as of the date this document was published, and may contain elements of subjectivity (including certain assumptions) or be based on incomplete information. There is no duty or obligation to update the contents of this document. This document is not intended to provide, and should not be relied on for accounting, legal, or tax advice, or investment recommendations. Pantera and its principals have made investments in some of the instruments discussed in this communication and may in the future make additional investments or trading decisions in connection with such instruments without further notice. This document solely reflects the opinion of the author, and does not reflect Pantera’s opinions.

 


References

[1] https://www.bcg.com/publications/2026/understanding-the-new-economics-of-ai-compute-markets

[2] https://seuc.senate.ca.gov/background-electricity-policy

[3] https://www.ferc.gov/electric-power-markets

[4] https://www.iso-ne.com/participate/support/faq/lmp

[5] https://www.eia.gov/todayinenergy/detail.php?id=3150

[6] https://www.ferc.gov/federal-statutes

[7] https://blogs.nvidia.com/blog/nvidia-ai-factory-compute/

[8] https://davefriedman.substack.com/p/who-benefits-from-compute-futures

[9] https://www.ice.com/evolution-of-brent-its-markets-and-why-its-ecosystem-is-relied-upon-by-commercial-participants

[10] https://x.com/0xfishylosopher/status/2090243620645515415?s=20

[11] https://sherwood.news/markets/the-ai-spending-boom-is-eating-the-us-economy/

Get the latest news in blockchain and crypto