A liquid futures market for AI compute would create something the field has never possessed: a continuously updated, financially consequential estimate of the future scarcity and economic value of machine computation. That sounds, at first, like a prediction market for AI progress. The idea is seductive. If frontier models require enormous quantities of specialized compute, and traders can price the future rental cost of that compute, perhaps the forward curve will reveal when the market expects another scaling wave, an agentic deployment boom, or a collapse in the cost of intelligence.
The literal version of this thesis is wrong. A compute futures contract is not an event contract paying out when an AI system crosses a capability threshold. Its price will mix expectations about GPU demand with chip supply, power constraints, hardware obsolescence, algorithmic efficiency, utilization, financing conditions, hedging pressure, benchmark construction, and risk premia. A higher futures price could signal faster expected AI progress, slower infrastructure delivery, a speculative risk premium, or simply a transformer shortage. A lower price could signal weak AI demand, or a major efficiency breakthrough that makes each GPU-hour more productive.
The stronger thesis is more interesting. A sufficiently rich compute-derivatives market could become a market-implied latent-state estimator for the AI economy. Its informational value would lie in the joint geometry of the curve rather than any headline price: maturities, chip-generation spreads, regional basis, training-versus-inference benchmarks, option-implied volatility, power-adjusted margins, and reactions to model releases, hardware announcements, export controls, and grid shocks. Read alongside capability evaluations, token prices, utilization data, and infrastructure indicators, the compute curve could become to AI what the yield curve is to macroeconomics: fallible, reflexive, contaminated by risk premia, yet too information-dense to ignore.
Compute futures are not prediction markets in the strict sense
A prediction market is normally designed around an outcome. A binary contract might pay one dollar if an event occurs and zero otherwise. Under suitable conditions, its price can be interpreted as an approximation of the market’s probability for that event. Other contract forms can elicit expectations about a future mean, median, or distribution.11
A compute future is different. It settles against the future price of a resource. Its terminal payoff is linked to a GPU-rental benchmark, not to a claim such as:
“A publicly available AI system will achieve a 50% task-completion horizon of eight hours by December 2027.”
The two market types are already touching, for what it’s worth. Polymarket cleared its first on-chain institutional block trade tied to AI compute in June 2026, settled against Ornn’s index of H100 rental pricing.12 A prediction-market venue clearing compute exposure is a sign of how porous the boundary may become. The conceptual distinction still stands.
It stands because commodity futures are not pure forecasts of future spot prices. A useful schematic decomposition is:
where is the futures price at time for maturity , is the market’s expectation of the future spot price, and is a maturity-dependent risk premium. The sign convention varies across models, but the substantive point does not: the quoted futures price contains both expectations and compensation for bearing risk.13
In physical commodities, the curve is also shaped by inventory, storage costs, financing, and the convenience yield of having immediate access to a scarce input. Empirical work shows that commodity risk premia vary across products and maturities and are related to hedging pressure and physical scarcity.14 Federal Reserve research finds that futures prices can be useful forecasting guides, while also documenting substantial deviations between spot and futures prices and the importance of storage and convenience yield.15
Compute complicates this structure further. A GPU-hour is not stored in the same way as copper or crude oil. Idle hardware can preserve future capacity, but unused time itself expires. Compute therefore resembles electricity, airline seats, freight capacity, and cloud spot instances more than a conventional storable commodity. It is capacity with a clock attached.
A compute curve would consequently be a market forecast only after several wedges are considered:
Calling the curve a prediction market for AI progress is therefore a metaphor. The useful question is whether it is a good metaphor, whether the market contains enough structured information to infer something about AI’s trajectory after these contaminating factors are modeled.
Why compute should contain information about progress
Compute is not synonymous with intelligence, but it has been one of the most measurable inputs to modern AI progress. OpenAI’s early analysis framed advances in AI as a function of three broad inputs, algorithms, data, and training compute, while emphasizing that compute is unusually quantifiable compared with the other two.16
Empirical scaling laws strengthened the relationship. Kaplan and colleagues found power-law relationships between language-model loss and model size, dataset size, and training compute over large ranges.17 DeepMind’s Chinchilla work later showed that model size and training data must be allocated together under a fixed compute budget; its central result was not merely that more compute helps, but that the way compute is allocated materially changes the capability obtained from it.18 Subsequent work has incorporated inference demand into compute-optimal scaling, showing that a model’s economically optimal training configuration depends on the quantity of downstream inference it is expected to serve.19
The causal chain is therefore real but mediated:
The chain also runs backward. Expected deployment revenue raises the value of training better models. That increases demand for training clusters, which raises demand for accelerators, memory, networking, power, cooling, and capital. The market price of compute is produced by both current capabilities and expectations about future ones.
This creates the possibility of forward-looking information. A laboratory that privately discovers a promising scaling regime may seek more future capacity. A cloud provider observing rapidly rising enterprise agent usage may revise its demand forecast. A semiconductor supplier may know that a production bottleneck will constrain deliveries. A power developer may see interconnection delays that prevent announced data centers from becoming operational. Each actor holds a different fragment of the AI system’s state. A liquid market can, in principle, force these fragments to meet.
The result would not be an oracle. It would be a compressed sufficient statistic for the marginal value of compute, conditional on the participants, benchmark, and market structure that exist at that moment.
The object being priced is not intelligence
A useful analytical model begins by separating physical compute from effective compute. Let:
where is installed physical accelerator capacity, is hardware performance per nominal accelerator-hour, is software and algorithmic efficiency, is utilization, and is a workload-specific quality adjustment covering memory, interconnect, reliability, and topology.
AI capability can then be represented schematically as:
where is data, is research and human capital, and captures architecture, objectives, post-training, tool use, and other algorithmic choices. Economic AI output depends on a related but distinct deployment function:
where includes product design, distribution, trust, regulation, organizational integration, and the availability of complementary human and physical systems. The spot price of compute is set elsewhere:
where represents energy and grid conditions, financing and capital constraints, and benchmark and market frictions.
These equations expose the central identification problem. The same movement in can be generated by very different changes in the underlying AI system. A rise in compute futures may reflect expected capability progress, because laboratories anticipate productive large runs. It may reflect expected deployment growth, because enterprises are consuming more inference. It may equally reflect hardware scarcity from delayed accelerator production, energy scarcity when projects cannot secure power, financing stress that raises the cost of building capacity, geopolitical segmentation as export controls divide regional markets, a larger risk premium when natural buyers grow desperate to hedge, or benchmark deterioration, where the contract ceases to match the capacity users actually need.
A fall is just as ambiguous. It may reflect weak AI demand or overbuilt data centers. It may also reflect a new accelerator generation, better model architectures, gains from quantization, batching, caching, or speculative decoding, migration to custom silicon, lower electricity costs, declining financing costs, or an increase in supply that enables more AI progress.
The sign is not invariant. Expensive compute can mean AI is accelerating. Cheap compute can also mean AI is accelerating.
The paradox of algorithmic efficiency
The most important reason a compute price cannot be read naively is that AI engineering constantly tries to destroy the scarcity being priced.
OpenAI defined algorithmic efficiency as the reduction in compute required to reach a fixed level of capability and documented rapid historical improvements in several machine-learning domains.20 Epoch AI’s current trends dashboard similarly separates advances in training compute, hardware, software efficiency, and inference prices. Its February 2026 update reports rapid improvement in pre-training compute efficiency and a steep decline in the price of inference at fixed performance, though the exact rates vary by task and methodology.21
An efficiency breakthrough creates two opposing effects.
The first is substitution. If a workload needs fewer GPU-hours, demand for raw compute falls: implies fewer GPU-hours per task, which pushes down. Under this channel, falling compute futures could be bullish for AI progress. The market is pricing a collapse in the resource cost of a given capability.
The second is rebound. Cheaper effective intelligence can create new use cases and increase total consumption: as the cost per useful task falls, the quantity of tasks demanded rises, and inference demand grows. If the elasticity of demand is high enough, total compute demand rises despite efficiency gains. This is a digital form of the Jevons paradox. Better inference kernels, smaller models, caching, and batching may reduce the cost per task while expanding the number, duration, and complexity of tasks sufficiently to raise total accelerator use.
Reasoning and agentic systems make this rebound especially plausible. A cheaper token may be consumed inside longer chains of thought, parallel sampling, verification loops, tool calls, or multi-agent workflows. The economically relevant unit becomes neither the token nor the GPU-hour, but the successful completion of a valuable task.
The implication for market interpretation is severe. A falling H100 curve accompanied by falling quality-adjusted token prices and rising task volume may indicate accelerating diffusion. The same falling curve accompanied by weak utilization and stagnant capability metrics may indicate a demand bust.
Training demand and inference demand tell different stories
A single compute index may combine two economically different markets.
Training compute is a wager on the frontier. Large training runs are lumpy, strategic, and concentrated. Their demand reflects expected scaling returns, competitive races among laboratories, research confidence, capital availability, access to data, and beliefs about the next architecture or post-training regime. Training demand is therefore closer to a venture-style option on future capability. A sudden rise in long-dated demand for tightly interconnected frontier-grade clusters could be informative about laboratories’ private expectations.
Inference compute is a wager on diffusion. Its demand reflects deployed usage: API traffic, coding agents, search and recommendation, enterprise automation, media generation, scientific workloads, robotics and control, consumer products. It is more closely connected to realized economic adoption than to frontier research alone.
The distinction is not cosmetic. Epoch AI estimates that the most prominent frontier laboratories did not use most of the world’s AI compute in 2025; substantial capacity was held or used by hyperscalers, other model developers, recommender systems, open-weight inference, and non-language workloads.22 A broad compute index could therefore rise because of mass deployment even if frontier scaling slows. Conversely, a few enormous training runs could tighten premium cluster capacity while leaving the broader rental market soft.
A prediction system should ideally maintain separate curves for tightly interconnected training clusters, latency-sensitive inference, throughput-oriented batch inference, memory-heavy long-context workloads, lower-cost commodity accelerators, and custom silicon where comparable pricing can be constructed. Market design is already moving in this direction: NATIVX’s COIL Index, the reference for ICE’s second announced compute-futures suite, publishes separate sub-indices for training, inference, graphics, and connectivity.23 Without segmentation of this kind, the market risks creating a beautifully liquid price for an economically incoherent average.
Compute is a multidimensional commodity
Standardization is the central market-design problem.
Silicon Data states that its indices cover major accelerator categories and draw on verified pricing records across countries, platforms, providers, regions, and lease structures. Its own methodology emphasizes that GPU pricing is multidimensional and that simple averages can conceal material heterogeneity.24 Earlier work on cloud spot markets similarly found that aggregation can stabilize an index, while individual virtual-machine prices remain volatile and operationally distinct.25
The heterogeneity is not a nuisance at the edge of the market. It is the substance of the product. An H100-hour differs according to SXM versus PCIe configuration, memory capacity and bandwidth, NVLink and cluster topology, the CPU, storage, and network it is paired with, utilization and contention, reliability and preemption risk, data residency and security, geography and latency, electricity price and carbon intensity, the software stack and orchestration, and whether the lease is reserved, on-demand, or interruptible.
Hardware generations compound the problem. A future B-series accelerator may deliver more useful work than an H100 while commanding a different price, power envelope, memory configuration, and software profile. A nominal hour is therefore a poor intertemporal unit unless it is adjusted for workload-specific output.
This suggests three possible market architectures.
The first is hardware-specific contracts, which settle on the rental price of a specific accelerator class. The approach is observable and relatively concrete, and it is where the announced U.S. products have started: Ornn’s index series reference named chips, from the H100 to the RTX 5090.26 Its weakness is rapid obsolescence and fragmentation.
The second is H100-equivalent or performance-adjusted contracts, which normalize different accelerators to a common performance unit. This creates continuity across generations, but equivalence depends on workload, precision, memory, interconnect, and software. NATIVX’s energy-normalized COIL Index is a variant of this direction: it prices compute per stable, auditable unit and deliberately trades hardware specificity for continuity, designed to sit alongside ICE’s natural gas and power complexes.27
The third is output-denominated contracts, which settle on the cost of producing a standardized inference or training output: quality-adjusted tokens, benchmark tasks, or a defined model workload. This comes closer to economic utility, but model quality, tokenizer choice, caching, latency, and provider behavior make output difficult to standardize. China’s reported exploration of AI-token futures illustrates this direction, contrasting with U.S. efforts centered on GPU-rental costs.28 A 2026 preprint has proposed a “standard inference token” and a contract design for token futures, but this remains a theoretical proposal rather than evidence that such a unit can sustain a durable market.29
The eventual predictive content will depend heavily on which architecture wins. A hardware-specific curve forecasts hardware scarcity. An output-denominated curve comes closer to forecasting the marginal price of useful machine cognition.
How to read the compute curve
The information would not live in the level alone. It would live in five dimensions.
The level
The price level estimates the current or future dollar cost of benchmarked capacity. A rising level may indicate a tightening supply-demand balance. It says little by itself about whether demand is coming from frontier training, inference adoption, non-AI workloads, or hedging pressure.
The slope
Let and denote near- and long-dated contracts. Backwardation, near prices above later prices, could indicate acute current scarcity expected to ease through new supply or obsolescence. Contango, later prices above near prices, could indicate expected demand growth, financing and carrying costs, future power constraints, or a positive term premium.
For compute, the slope may also encode expected generational replacement. A steep decline in an H100 curve can coexist with a rising B200 curve if the market expects rapid migration toward newer hardware.
Inter-generation spreads
Define:
where quality-adjusts for expected performance. This spread may reveal the pace of hardware adoption, scarcity of new-generation systems, expected software readiness, the depreciation rate of older accelerators, and the market value of memory and interconnect improvements. A rapidly widening spread could indicate that frontier workloads increasingly require the new generation. A narrowing spread could indicate supply normalization, software bottlenecks, or strong demand for cheaper inference on older chips.
Regional basis
Let:
Regional basis could price power availability, grid congestion, data sovereignty, export controls, latency, political risk, cooling and water constraints, and local financing conditions. Because data centers are spatially concentrated, power bottlenecks can be locally severe even when global electricity shares remain modest. The International Energy Agency projects global data-center electricity consumption to reach roughly 945 TWh by 2030 in its base case and stresses that localized concentration makes grid integration especially challenging.30
A rising regional compute basis may therefore be less a forecast of AI capability than a forecast of infrastructure failure.
Implied volatility and skew
Options on compute futures would reveal the price of tail risk. High upside skew could indicate fear of shortages, model-driven demand shocks, export controls, or power failures. High downside skew could indicate overbuilding, obsolescence, demand disappointment, or an efficiency shock. Calendar spreads in implied volatility could reveal when the market expects uncertainty to resolve.
For forecasting AI, options may be more informative than futures. A futures level compresses beliefs into an average price. Options reveal the market’s pricing of discontinuities, the very object of interest when discussing sudden capability jumps or infrastructure shocks.
A regime map: the same price movement can imply opposite futures
| Market observation | Plausible structural cause | Interpretation for AI progress |
|---|---|---|
| Long-dated compute rises; chip-delivery expectations unchanged | Stronger expected training or inference demand | Potentially bullish |
| Long-dated compute rises; power basis rises simultaneously | Grid and energy constraints | Ambiguous or bearish for realized progress |
| H100 falls; B200 rises | Generational migration | Bullish for frontier capability, bearish for old hardware |
| All hardware curves fall; quality-adjusted inference prices fall faster; usage rises | Efficiency and diffusion | Strongly bullish |
| All curves fall; utilization and AI revenue fall | Demand destruction or overbuild | Bearish |
| Near curve spikes; long curve stable | Temporary outage or delivery shock | Little information about long-run progress |
| Long curve rises; option upside skew rises after a scaling result | Market reprices future frontier demand | Potential leading signal |
| Compute rises; capability metrics stagnate | Input escalation without output improvement | Evidence of declining marginal returns |
| Compute flat; capability metrics rise rapidly | Algorithmic progress dominates physical scaling | Bullish, with compute prices understating progress |
| Regional basis widens under export controls | Market fragmentation | Global progress may continue while access becomes unequal |
This table captures the central rule: a compute price is interpretable only conditionally.
The compute curve as a latent-state estimator
The correct analogy is not a binary prediction market. It is a state-estimation problem.
Suppose the hidden state of the AI economy is:
The observable market vector is:
A dynamic factor or state-space model could represent:
Capability and economic outcomes would then be modeled as:
where might include frontier benchmark improvement, loss at fixed compute, task-completion time horizon, quality-adjusted inference cost, AI revenue or usage, observed training-run scale, and energy consumed by accelerated servers.
The research question is not whether the compute price equals expected AI progress. It is:
Does the market-derived state vector improve out-of-sample forecasts of AI capability or adoption after controlling for public technical and infrastructure information?
That is empirically testable.
What should count as AI progress?
A prediction signal is meaningless without a target variable. AI progress is multidimensional, and at least four targets should be separated.
The first is scientific or benchmark capability: loss, reasoning accuracy, coding performance, cyber capability, scientific problem-solving, or multimodal competence. These measures are relatively easy to update but vulnerable to saturation, contamination, and benchmark-specific optimization.
The second is autonomous task horizon. METR has proposed measuring the length of tasks that AI agents can complete at a specified reliability. Its work reports an approximately exponential historical trend, while emphasizing uncertainty in extrapolating it to future real-world autonomy.31 This kind of metric is attractive because it translates model capability into the duration and complexity of coherent action.
The third is quality-adjusted cost. Progress can mean producing the same capability more cheaply:
This target is essential because a model that matches yesterday’s frontier at one-tenth the cost may matter more economically than a marginal benchmark leader.
The fourth is realized productivity and adoption. Capability does not automatically become economic output. METR’s randomized study of experienced open-source developers found that early-2025 AI tools made participants slower in the tested setting, illustrating the gap that can exist between benchmark excitement and realized productivity.32 That result is a snapshot, not a universal verdict on AI, but it demonstrates why deployment must be measured separately.
A compute curve might forecast these targets with different signs and horizons. Training-cluster futures may lead frontier capability. Commodity inference prices may lead diffusion. Regional compute-power spreads may lead infrastructure deployment. None should be treated as a universal “AI progress index.”
Event studies that could reveal informational content
The first years of compute futures would provide a natural laboratory. Researchers should examine how the curve responds to distinct classes of news.
Model and algorithm events include the release of a model with unexpectedly strong capabilities, evidence that test-time compute improves performance, a major distillation or quantization result, breakthroughs in sparse architectures, memory, or routing, and evidence of diminishing returns to scaling. Hardware events include accelerator announcements, production delays, HBM shortages, interconnect improvements, custom-silicon launches, and changes in export controls. Infrastructure events include grid interconnection approvals or cancellations, power-purchase agreements, turbine and transformer delivery shocks, major data-center financing, and water or permitting restrictions. Demand events include enterprise adoption reports, agent usage, cloud utilization, API price changes, large customer contracts, and evidence of AI-driven productivity.
For each event class, one could estimate abnormal movements in curve level, slope, generation spreads, regional basis, implied volatility, and volume and open interest.
A genuinely informative market should respond differently to an algorithmic efficiency shock than to a power-supply shock. If every headline merely moves one generic “AI hype factor,” the market will have little analytical value.
Can the curve lead capability releases?
The strongest prediction-market thesis would be supported if compute prices systematically move before public capability evidence. That could happen through several channels.
Private procurement is the first. Labs reserve capacity before training begins. Suppliers, clouds, brokers, lenders, and infrastructure firms may observe demand without knowing the model details.
Financing is the second. A laboratory seeking capital or signing long-duration contracts reveals confidence through costly commitments. Futures participants may infer the expected value of upcoming runs. The scale of such commitments is already extraordinary: Google agreed in mid-2026 to pay SpaceX roughly $920 million per month for compute capacity from October 2026 through June 2029.33
Supply-chain telemetry is the third. Accelerator orders, HBM allocation, networking demand, and power contracting may reveal the scale of projects before model releases.
Informed hedging is the fourth. An AI company expecting a demand surge may buy protection against rising inference costs. Its hedge can move prices even if it never publicly explains the reason.
A key empirical test would compare compute-market moves with later model launches, benchmark jumps, task-horizon improvements, API demand, revenue, and cluster completion. If the curve contains non-public but lawful commercial information, it may become a leading indicator. If it simply reacts to press releases, it will be a coincident sentiment index.
When the compute curve will lie
Even a liquid market can produce a misleading narrative. Eight failure modes deserve advance attention.
Risk premia can masquerade as expectations. Natural buyers may be more desperate to hedge than natural sellers. A positive insurance premium can push futures above expected spot prices, and interpreting the quote as a literal forecast would overstate expected scarcity.
Liquidity can be thin. A new contract may have low volume, wide bid-ask spreads, concentrated positions, and fragile market making. Prices can be moved by balance-sheet constraints rather than information.
Benchmark basis can drift. The settlement index may not match the capacity a user needs. A benchmark can be statistically broad yet economically wrong for frontier training, secure enterprise inference, or a particular region.
Adverse selection can hollow out the price. Participants with the best information may avoid the public market if trading reveals their plans, hedging instead through bilateral contracts, vertical integration, power agreements, or cloud reservations. The largest compute commitments today, including deals on the scale of the Google-SpaceX agreement, are struck privately and may never touch an exchange.34
Manipulation is possible. An actor with a large physical position may benefit from influencing the benchmark or futures settlement. Strong index governance, transaction verification, position limits, and surveillance will be essential.
Obsolescence is endogenous. The expected useful life of the benchmarked hardware may change faster than the contract maturity. A technically valid settlement price can become irrelevant to the frontier.
Narratives can correlate everything. Compute futures, semiconductor equities, and AI-company valuations may all respond to the same speculative factor. Apparent forecasting power may disappear after controlling for broader AI sentiment.
Capability and value can decouple. A system can become more capable without becoming more profitable, or more profitable without advancing the frontier. Markets care about cash flows and scarcity, not philosophical intelligence.
Reflexivity: the market will change the future it predicts
A compute curve would not merely observe AI progress. It would influence it.
A high long-dated price can improve the expected revenue of compute suppliers. That can support data-center project finance, GPU-backed lending, power investment, long-term procurement, new cloud entrants, and accelerator production. The feedback loop is:
The inverse loop can cancel projects and reduce future supply.
The financial superstructure is assembling before the underlying market exists. ETF sponsors filed for leveraged and inverse products tied to compute futures within days of the CME announcement, and Silicon Data’s rental benchmarks have already appeared in corporate disclosures, including SpaceX’s IPO prospectus.35 Speculative capital and reference-pricing effects will be present from the first day of trading, not phased in gradually.
This reflexivity makes the compute curve closer to a control signal than a passive thermometer. It can coordinate the construction of the system whose future price it quotes. That is economically useful, but it complicates causal interpretation. If a high price predicts abundant future compute because it caused investment, the initial “forecast error” may actually demonstrate successful coordination.
AI supercomputers already resemble industrial megaprojects. A dataset of more than 500 systems found that from 2019 to 2025 their computational performance doubled roughly every nine months, while estimated hardware acquisition cost and power requirements doubled roughly annually.36 When projects reach this scale, financing conditions become part of the capability-production function.
The compute curve may therefore forecast AI progress partly by financing it.
The yield-curve analogy, and its limits
The best analogy is the yield curve.
The yield curve is not a pure forecast of future short-term interest rates. It also contains term premia, liquidity effects, regulation, central-bank policy, collateral demand, and risk appetite. Yet economists still study its level, slope, curvature, inversions, and cross-market relationships because the curve aggregates consequential beliefs and constraints.
A mature compute curve could play a similar role. Its level would measure current scarcity of benchmarked compute; its slope, the expected evolution of scarcity and investment; its curvature, transitions between hardware generations or infrastructure phases. Regional basis would price geopolitical and energy fragmentation, generation spreads the rate of technological obsolescence, implied volatility the uncertainty and tail risk, compute-power margins the profitability of converting electricity into machine computation, and token-compute spreads the state of software efficiency and product-layer value capture.
The analogy breaks at one decisive point. Interest rates are denominated in a stable financial unit. Compute quality is endogenous to technological progress. The underlying itself mutates. The market is attempting to construct a yield curve over a unit whose productive meaning changes every few months.
For that reason, the most informative object may not be a single curve but a surface:
AI progress would be inferred from the deformation of this surface through time.
A proposed Compute-Implied AI Progress Index
A serious index should avoid the mistake of equating expensive compute with rapid progress. One possible framework would combine five standardized factors.
The first factor is frontier scarcity, measured through prices for tightly interconnected, current-generation training clusters. The second is deployment demand, measured through inference-oriented capacity, utilization, API volume, and output-denominated prices. The third is efficiency, measured through the spread between raw compute prices and quality-adjusted output prices:
If raw compute remains expensive while useful output becomes dramatically cheaper, efficiency is improving. The fourth factor is infrastructure constraint, measured through regional compute basis, electricity curves, power interconnection, and compute-power margins. The fifth is innovation option value, measured through long-dated implied volatility and generation spreads around known research and hardware milestones.
The combined index would be estimated rather than hand-weighted. A dynamic model could be trained to forecast multiple progress targets and constrained to preserve interpretability. Its output should be a distribution, not a point estimate. For example:
This still would not be a literal prediction-market probability. It would be a model-implied forecast using market data as inputs. The distinction should be maintained in every public presentation.
Governance: a public early-warning layer for AI
Compute has attracted governance interest because it is relatively quantifiable, detectable, excludable, and produced through a concentrated supply chain. Researchers have argued that compute can support visibility into AI development, resource allocation, and enforcement, while warning that poorly designed controls could create privacy, economic, and centralization risks.37
A regulated compute market could add a new kind of visibility. Unusual movements in long-dated frontier-grade contracts, upside option skew, regional basis under export controls, physical-financial divergence, or the concentration of positions might provide early warnings of emerging shortages, major training demand, financial instability, or geopolitical segmentation. Regulators are already engaged with the product’s foundations: contract specifications, settlement procedures, and benchmark construction are expected to face scrutiny from the Commodity Futures Trading Commission before launch.38
The signal would be indirect. Exchanges would see positions, not model weights or training objectives. Most compute demand would remain benign and commercially sensitive. Surveillance could easily become overbroad. Still, the market could supply aggregate information that neither voluntary laboratory disclosures nor public benchmarks currently provide.
The governance opportunity is therefore narrower than letting the market regulate AI:
Use market telemetry as one input in a plural monitoring system that also includes capability evaluations, chip and cluster data, energy infrastructure, corporate reporting, and technical audits.
Such a system could be useful precisely because traders have money at stake. Cheap talk about AI timelines is abundant. A hedged procurement position is a costly signal.
An empirical research program
The thesis should be treated as a research program, not a slogan.
The dataset would collect tick or daily compute-futures prices; open interest, volume, and trader categories; benchmark composition and revisions; spot rental transactions; cloud reservation prices; chip shipments and delivery times; power prices and interconnection data; data-center completions; model releases and capability evaluations; quality-adjusted inference prices; and AI usage and revenue indicators.
Six hypotheses follow.
First, compute curves lead capability releases. Test whether curve innovations predict benchmark or task-horizon gains after controlling for public announcements, semiconductor equities, and broad risk appetite.
Second, generation spreads predict architectural migration. Test whether H100-B200 or equivalent spreads lead changes in workload allocation and frontier-cluster composition.
Third, output-input spreads measure algorithmic efficiency. Test whether movements between token or task prices and raw GPU prices predict independent estimates of software efficiency.
Fourth, regional basis predicts infrastructure delivery. Test whether compute basis predicts project delays, power constraints, and local utilization.
Fifth, options contain discontinuity information. Test whether implied skew rises before major capability or supply shocks and whether it outperforms analyst forecasts.
Sixth, financialization changes investment. Use exchange launches, contract introductions, or liquidity thresholds as quasi-experiments to estimate whether hedgeability lowers financing costs or accelerates infrastructure construction.
Every model should be evaluated out of sample against simple baselines: the last observation, a linear trend, a semiconductor equity index, announced capex, public chip-shipment forecasts, capability trend extrapolation, and expert surveys. The burden of proof is straightforward: market data must add predictive information beyond what was already public.
What could ultimately be written on the tape?
If the market succeeds, commentators will be tempted to narrate every price movement as a referendum on artificial general intelligence. That would be intellectually lazy.
A better interpretation hierarchy is:
- First: What changed in the physical supply-demand balance for benchmarked compute?
- Second: Was the movement driven by training, inference, power, hardware, financing, or risk premia?
- Third: What does that structural change imply for effective compute?
- Fourth: What does effective compute imply for capability, deployment, or economic output?
- Finally: Does the move survive comparison with independent evidence?
Only after those steps should anyone claim that “the market is pricing faster AI progress.”
The compute curve will be most informative when it disagrees with the dominant narrative. AI equities might rise while long-dated compute demand falls. Raw GPU prices might fall while quality-adjusted inference costs collapse and usage explodes. Frontier-grade capacity might tighten while broad commodity inference remains cheap. Power-adjusted compute margins might deteriorate despite bullish capability releases. Option markets might price a large upside tail while consensus forecasts remain smooth. These divergences can reveal whether the bottleneck lies in intelligence, infrastructure, diffusion, or finance.
Conclusion: a market in the shadow of intelligence
Compute futures will not directly price intelligence. They will price claims on machines, electricity, memory, networks, locations, contract terms, and time. The resulting curve will be noisy, strategic, and distorted by the same forces that complicate every commodity market.
Yet that does not make it trivial.
Modern AI is produced through a physical-economic pipeline. Research ideas are translated into training runs; training runs into capabilities; capabilities into products; products into inference demand; demand into data centers, power plants, chip orders, and financing. Compute is one of the few variables that touches almost every stage.
A futures market would place a public price on the marginal capacity flowing through that pipeline. Its movements could aggregate knowledge held by laboratories, clouds, hardware firms, utilities, data-center developers, lenders, and users. That makes it a potentially powerful observational instrument.
The right claim is therefore neither:
“Compute futures will predict AGI,”
nor:
“Compute prices tell us nothing about AI progress.”
The defensible claim is:
A mature compute-derivatives complex could become the first continuously traded, financially disciplined estimate of the AI economy’s hidden state.
Its predictive power would come from structure, not mysticism: the curve, the spreads, the basis, the volatility surface, and their relationship to independent measures of capability and adoption.
The yield curve does not tell us the future. It tells us what a consequential market is collectively positioned for, after risk and constraint have left their fingerprints on the price. A compute curve may do the same for AI.
And because that curve will influence financing, construction, access, and strategy, it will not merely predict the trajectory of machine intelligence. It will become one of the mechanisms through which that trajectory is made.
Sources and cross-references
Source-status note
This article was source-checked on July 11, 2026. The CME-Silicon Data, ICE-Ornn, and ICE-NATIVX products were publicly announced and described as subject to regulatory review or approval; none had begun trading. The essay therefore refers to them as planned or developing markets, not as mature liquid markets with established forecasting records. Reports of bank participation and ETF filings describe intentions and proposals, not operating businesses. Claims about the future informational value of these markets are analytical hypotheses and proposed research questions unless explicitly supported by historical evidence from other futures or prediction markets.
Milken Institute, “A Conversation with BlackRock CEO Larry Fink and Brookfield Corporation CEO Bruce Flatt,” May 5, 2026, event page and transcript access. The original article supplied for this essay is Andreja Stojanovic, “BlackRock says ‘compute futures’ could become a new asset class,” Finbold, May 8, 2026, link.↩︎
CME Group, “CME Group and Silicon Data Partner to Launch First Compute Futures,” May 12, 2026, link. CME’s product information page is Compute Futures.↩︎
Intercontinental Exchange, “ICE and Ornn to Launch GPU Compute Futures Contracts,” May 19, 2026, link. See also Ornn, “The OCPI: The Reference Price for Compute,” link.↩︎
Blockspace, “ICE and Ornn to launch GPU compute futures contracts,” May 19, 2026, link. Notes that OCPI is built only from printed transactions, that its series cover the H100, H200, B200, and RTX 5090, that Ornn executed the first compute swap using cleared prices in December 2025, and that ICE joins a small but growing group of venues targeting compute derivatives.↩︎
PYMNTS, “Big Banks Eye New AI Compute Trading Market,” June 9, 2026, link, summarizing reporting by The Information on Goldman Sachs and JPMorgan’s early-stage exploration of compute trading. The same report covers Polymarket’s first on-chain institutional block trade settled against Ornn’s H100 rental index (June 2, 2026) and Google’s disclosed agreement to pay SpaceX $920 million per month for compute capacity from October 2026 through June 2029. The banks’ exploration was described as preliminary and may not proceed.↩︎
CNBC, “The new oil? Inside the effort to turn AI computing power into a tradeable commodity,” June 16, 2026, link. Reports the ProShares and Rex Shares ETF filings (including leveraged and inverse products), SpaceX’s reference to Silicon Data rental-rate data in its IPO prospectus, and expectations of CFTC scrutiny of contract specifications, settlement, and benchmark construction.↩︎
Intercontinental Exchange and NATIVX, “ICE and NATIVX to Launch Energy-Normalized Compute Futures,” July 2026, Business Wire release, reproduction. The COIL Index tracks tokenized, energy-normalized compute and connectivity and publishes COIL-T (training), COIL-I (inference), COIL-G (graphics), and COIL-CO (connectivity) sub-indices. Contracts are planned as U.S.-dollar-denominated and cash-settled, subject to regulatory approval.↩︎
CNBC, “Traders will soon be able to bet on computer chip prices as AI drives costs skyward,” May 12, 2026, link. Recounts the precedent of Enron’s broadband division attempting to sell unused fiber-optic capacity in the late 1990s before the company’s failure.↩︎
F. A. Hayek, “The Use of Knowledge in Society,” American Economic Review 35, no. 4 (1945), hosted by Econlib, link.↩︎
Justin Wolfers and Eric Zitzewitz, “Prediction Markets,” Journal of Economic Perspectives 18, no. 2 (2004): 107–126, link.↩︎
Justin Wolfers and Eric Zitzewitz, “Prediction Markets,” Journal of Economic Perspectives 18, no. 2 (2004): 107–126, link.↩︎
PYMNTS, “Big Banks Eye New AI Compute Trading Market,” June 9, 2026, link, summarizing reporting by The Information on Goldman Sachs and JPMorgan’s early-stage exploration of compute trading. The same report covers Polymarket’s first on-chain institutional block trade settled against Ornn’s H100 rental index (June 2, 2026) and Google’s disclosed agreement to pay SpaceX $920 million per month for compute capacity from October 2026 through June 2029. The banks’ exploration was described as preliminary and may not proceed.↩︎
Jonathan Hambur and Nick Stenner, “The Term Structure of Commodity Risk Premiums and the Role of Hedging,” Reserve Bank of Australia Bulletin, March 2016, link.↩︎
Jonathan Hambur and Nick Stenner, “The Term Structure of Commodity Risk Premiums and the Role of Hedging,” Reserve Bank of Australia Bulletin, March 2016, link.↩︎
Trevor A. Reeve and Robert J. Vigfusson, “Evaluating the Forecasting Performance of Commodity Futures Prices,” Board of Governors of the Federal Reserve System, International Finance Discussion Paper No. 1025, 2011, link.↩︎
Dario Amodei and Danny Hernandez, “AI and Compute,” OpenAI, May 16, 2018, link.↩︎
Jared Kaplan et al., “Scaling Laws for Neural Language Models,” arXiv:2001.08361, 2020, link.↩︎
Jordan Hoffmann et al., “Training Compute-Optimal Large Language Models,” arXiv:2203.15556, 2022, link.↩︎
Alexander Sardana et al., “Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws,” arXiv:2401.00448, 2024, HTML.↩︎
Danny Hernandez and Tom Brown, “Measuring the Algorithmic Efficiency of Neural Networks,” OpenAI, May 5, 2020, overview.↩︎
Epoch AI, “Trends in Artificial Intelligence,” updated February 5, 2026, link. The dashboard reports separate estimates for training compute, hardware, software efficiency, inference prices, investment, and data-center scale; estimates should be read with the uncertainty and methodology provided on the page.↩︎
Josh You, “How Much AI Compute Do Frontier Labs Use?” Epoch AI, May 20, 2026, link.↩︎
Intercontinental Exchange and NATIVX, “ICE and NATIVX to Launch Energy-Normalized Compute Futures,” July 2026, Business Wire release, reproduction. The COIL Index tracks tokenized, energy-normalized compute and connectivity and publishes COIL-T (training), COIL-I (inference), COIL-G (graphics), and COIL-CO (connectivity) sub-indices. Contracts are planned as U.S.-dollar-denominated and cash-settled, subject to regulatory approval.↩︎
Carmen Li, “Building a Robust GPU Index,” Silicon Data, March 12, 2026, link. For background on the initial H100 rental index, see Samuel K. Moore, “Silicon Data Launches First GPU Price Index,” IEEE Spectrum, May 27, 2025, link.↩︎
Supreeth Shastri and David Irwin, “Cloud Index Tracking: Enabling Predictable Costs in Cloud Spot Markets,” arXiv:1809.03110, 2018, link.↩︎
Blockspace, “ICE and Ornn to launch GPU compute futures contracts,” May 19, 2026, link. Notes that OCPI is built only from printed transactions, that its series cover the H100, H200, B200, and RTX 5090, that Ornn executed the first compute swap using cleared prices in December 2025, and that ICE joins a small but growing group of venues targeting compute derivatives.↩︎
Intercontinental Exchange and NATIVX, “ICE and NATIVX to Launch Energy-Normalized Compute Futures,” July 2026, Business Wire release, reproduction. The COIL Index tracks tokenized, energy-normalized compute and connectivity and publishes COIL-T (training), COIL-I (inference), COIL-G (graphics), and COIL-CO (connectivity) sub-indices. Contracts are planned as U.S.-dollar-denominated and cash-settled, subject to regulatory approval.↩︎
Reuters, “China works on AI token futures market, sources say, in race with US,” May 28, 2026, link. Reuters reported that the design work was preliminary and that no launch timeline had been established.↩︎
Yicai Xing, “AI Token Futures Market: Commoditization of Compute and Derivatives Contract Design,” arXiv:2603.21690, 2026, link. This is a preprint proposing a market design and simulation, not validation from an operating token-futures market.↩︎
International Energy Agency, “Energy Demand from AI,” in Energy and AI, 2025, link. See also the executive summary.↩︎
METR, “Measuring AI Ability to Complete Long Tasks,” March 19, 2025, with updated time-horizon data, link. METR explicitly discusses methodological uncertainty and warns that trend extrapolation is not a guarantee.↩︎
Joel Becker, Nate Rush, Beth Barnes, and David Rein, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity,” METR, July 10, 2025, link. The result applies to the study’s specified tools, participants, tasks, and period.↩︎
PYMNTS, “Big Banks Eye New AI Compute Trading Market,” June 9, 2026, link, summarizing reporting by The Information on Goldman Sachs and JPMorgan’s early-stage exploration of compute trading. The same report covers Polymarket’s first on-chain institutional block trade settled against Ornn’s H100 rental index (June 2, 2026) and Google’s disclosed agreement to pay SpaceX $920 million per month for compute capacity from October 2026 through June 2029. The banks’ exploration was described as preliminary and may not proceed.↩︎
PYMNTS, “Big Banks Eye New AI Compute Trading Market,” June 9, 2026, link, summarizing reporting by The Information on Goldman Sachs and JPMorgan’s early-stage exploration of compute trading. The same report covers Polymarket’s first on-chain institutional block trade settled against Ornn’s H100 rental index (June 2, 2026) and Google’s disclosed agreement to pay SpaceX $920 million per month for compute capacity from October 2026 through June 2029. The banks’ exploration was described as preliminary and may not proceed.↩︎
CNBC, “The new oil? Inside the effort to turn AI computing power into a tradeable commodity,” June 16, 2026, link. Reports the ProShares and Rex Shares ETF filings (including leveraged and inverse products), SpaceX’s reference to Silicon Data rental-rate data in its IPO prospectus, and expectations of CFTC scrutiny of contract specifications, settlement, and benchmark construction.↩︎
Konstantin F. Pilz, James Sanders, Robi Rahman, and Lennart Heim, “Trends in AI Supercomputers,” arXiv:2504.16026, 2025, HTML. The authors note dataset coverage and estimation limitations.↩︎
Girish Sastry et al., “Computing Power and the Governance of Artificial Intelligence,” arXiv:2402.08797, 2024, HTML.↩︎
CNBC, “The new oil? Inside the effort to turn AI computing power into a tradeable commodity,” June 16, 2026, link. Reports the ProShares and Rex Shares ETF filings (including leveraged and inverse products), SpaceX’s reference to Silicon Data rental-rate data in its IPO prospectus, and expectations of CFTC scrutiny of contract specifications, settlement, and benchmark construction.↩︎