Entrar / Cadastrar

As inference gets cheaper, what can NVIDIA earn? / Investigação da empresa

Will cheaper AI inference weaken NVIDIA or expand its market?

NVIDIA’s operating profit is still growing; Amazon says its own Trainium chips already run most Bedrock inference. More inference use does not mean every new order goes to NVIDIA.

Redação DisconfirmAI · Primeira publicação · · Última atualização substancial ·

A análise atual

Through FY2027 Q2, NVIDIA was still earning more: revenue rose 17.9% sequentially, operating income 19.0%, and gross margin remained near 75%. Google reports lower costs per AI Mode response; that service-specific efficiency gain can coexist with growth in NVIDIA’s company-wide earnings. The more concrete pressure is customer hardware choice: Amazon says Trainium runs most Bedrock inference. The disclosures do not establish lower costs as the cause of growth, but in-house chips already handle production inference, making the competition more than a future threat.

01

Earnings are growing. So is the competition.

NVIDIA’s operating income for the quarter ended July 26, 2026 was $63.734 billion, up from $53.536 billion in the previous quarter. Meanwhile, Amazon says its own Trainium chips already run most Bedrock inference. NVIDIA’s earnings growth and customers’ use of alternative chips are happening together.

In that quarter, NVIDIA’s company revenue rose 17.9% sequentially and operating income 19.0%. Gross margin stayed near 75%, close to the previous quarter. Revenue, gross profit and operating income all increased across five quarters. The company statements do not yet show cheaper computation squeezing operating earnings. They cover more than inference and do not measure GPU units.

Revenue and operating profit rise across five quarters

USD billions

Bilhões USD · trimestres fiscais individuais

  • Revenue
  • Gross profit
  • Operating income
Revenue and operating profit rise across five quarters · Bilhões USD · trimestres fiscais individuaisFive NVIDIA company quarters in USD billions, not inference revenue or GPU units. Displayed values = raw USD / one billion. FMP MCP records acquired October 9, 2026; main-quarter disclosure August 26. FY2026 Q2 2025-07-27; FY2026 Q3 2025-10-26; FY2026 Q4 2026-01-25; FY2027 Q1 2026-04-26; FY2027 Q2 2026-07-26025507510046.7433.8528.4457.0141.8536.0168.1351.0944.381.6261.1653.5496.2272.1463.73
FY2026 Q2
FY2026 Q3
FY2026 Q4
FY2027 Q1
FY2027 Q2

Consultado em: 2026-10-09 · FMP ↗

Ver os valores originais ↓
Trimestre fiscal / encerramentoRevenue · USDGross profit · USDOperating income · USD
FY2026 Q2
2025-07-27
46,743,000,00033,853,000,00028,440,000,000
FY2026 Q3
2025-10-26
57,006,000,00041,849,000,00036,010,000,000
FY2026 Q4
2026-01-25
68,127,000,00051,093,000,00044,299,000,000
FY2027 Q1
2026-04-26
81,615,000,00061,157,000,00053,536,000,000
FY2027 Q2
2026-07-26
96,221,000,00072,142,000,00063,734,000,000
Five NVIDIA company quarters in USD billions, not inference revenue or GPU units. Displayed values = raw USD / one billion. FMP MCP records acquired October 9, 2026; main-quarter disclosure August 26.

Lower computation costs can make previously uneconomic AI features viable or fund deeper reasoning within an existing budget. How many orders NVIDIA gains depends on how much computation those uses need and which systems customers choose.

The same price cut can affect bills, resources and chip revenue differently.

02

Unit prices, total bills and chip earnings are different accounts

Within a consistent billing boundary, a customer’s total spending is average paid price times paid quantity. If quantity growth more than offsets the price reduction, the bill can still rise. Longer contexts, additional answer candidates and verification can also consume more computation per job. A cheaper unit does not necessarily mean a smaller service budget.

API prices are not equipment requirements. Smaller models, more efficient computation, cheaper hardware, better utilization and provider concessions can all reduce the bill. Some change the resources needed for the same work; a concession may not. Compare the full resources required to meet the same quality standard, including failures and retries.

The calculator below illustrates resources only. Assume each task uses 40% less and task count grows 60%: the resource index becomes 96, still 4% below the starting level. Offsetting that saving requires about 66.7% more tasks. These are adjustable assumptions, not an API discount or measured demand growth.

How much task growth offsets a 40% resource saving?

Diagrama e comparação de dados

Adjust the assumptions to see how three conditions combine

NVIDIA resource-demand index · baseline 100

96.0

1.60 × 0.60 × 1.00 × 100 = 96.0

Under these assumptions, task volume must rise 66.7% to maintain baseline resource demand.

Arithmetic example with fixed task quality, system boundary and period. Resource efficiency is different from API pricing. This does not estimate actual elasticity, GPU purchases, revenue or stock prices.

Assume 40% fewer resources per task, 60% more tasks and a relative NVIDIA-system usage factor of 100: resource index 96; offsetting the saving requires about 66.7% task growth. Quality and task mix stay fixed; failed attempts count. The factor of 100 means unchanged from baseline, not market share. Not API pricing, measured demand or a revenue forecast.

NVIDIA’s revenue also depends on the customer’s hardware choice. Additional computation may run on GPUs, TPUs or Trainium, with different product prices and margins. More calls, more computation and higher NVIDIA earnings are not interchangeable results.

Google’s paid usage shows why computation demand cannot be read from list prices alone.

03

Google’s API processing rate rises as Search response costs fall

In its July 22, 2026 Q2 CEO remarks, Google reported model API processing at about 22 billion tokens per minute, compared with roughly 16 billion in the previous disclosure. The rounded figures imply an increase of about 37.5%, and Google still described supply constraints. These are rates at two disclosure points, not cumulative quarterly usage.

What usage did Google report?

Diagrama e comparação de dados
ServiceReported dataMeaning
Model APIsAbout 16bn → 22bn tokens/minuteDisclosed rate about +37.5%; supply still constrained
AI Mode and SearchLowest response cost; incremental Search queriesSeparate measures: efficiency and use grow together
Paid Cloud customersNearly 500, each >1tn tokens over 12 monthsCommercial use has scale
Google CEO remarks, July 22, 2026. API figures are rates at two disclosure points; 37.5% = 22/16−1. AI Mode cost and Search queries describe another service. Paid customer tokens cover twelve months. These rows do not form an elasticity estimate or GPU procurement measure.

In a separate service, Google said AI Mode response costs were the lowest since launch while overall Search gained incremental queries. API and Search measures cannot form a single price–demand curve. They nevertheless show an operator reporting improved efficiency alongside expanding use, rather than an inevitable contraction as costs fall.

Commercial use extends beyond trials. Google also reported nearly 500 Cloud customers each processing more than one trillion paid tokens over the preceding twelve months. Token counts vary with model and workload mix and cannot be converted into GPU purchases. They remain closer to actual demand than product launches, downloads or financing announcements.

Deeper reasoning can consume some of the budget that efficiency releases. OpenAI’s 2024 o1 research showed improved performance on some reasoning tasks with more computation when generating answers. Customers can spend less on the same work or use the same money to seek better answers. Existing paid usage supports expansion more convincingly than price cuts alone support a smaller compute market.

Google reports growing use; Amazon reports a different development: Trainium already runs most Bedrock inference.

04

Trainium is in production, including for frontier models

Amazon says Trainium already runs most Bedrock inference. It also says Claude models use more than one million Trainium2 chips across training and operation. The first statement concerns Amazon’s own platform; the second spans two kinds of computation. Neither measures global inference share, but both describe deployment rather than an unfulfilled roadmap.

Which buyers can make dedicated chips work?

Diagrama e comparação de dados
Buyer and workRoutes to compareBasis
Scale, sustained work, in-house optimizationDedicated chips and in-house platformsAmazon says most Bedrock inference runs on Trainium
Changing models, mixed work, rapid deploymentReusable GPU software and network systemsAdaptation and maintenance affect deployment costs
Large clouds with varied servicesMultiple hardware routesGoogle offers own and NVIDIA accelerators with portability
Bedrock deployment is Amazon’s operating statement; the other rows explain selection conditions. The Claude statement about more than one million Trainium2 chips spans training and operation. No global share or matched cost winner is shown.

Claude challenges the split between dedicated chips for simple models and GPUs for the smartest ones. Whether customers can make dedicated chips work depends on request volume, sustained workloads, and their ability to optimize models, compilers and chips together. At sufficient scale, long-term engineering work can be spread across more services. AWS has those organizational capabilities and an operating deployment.

For NVIDIA, the specific competition is that such buyers can route more new work to their own chips. Reusable GPU tools still matter to customers with changing models, mixed workloads and urgent deployment needs. Less low-level adaptation may be worth more than a cheaper individual chip. Engineering and service costs divide these settings, not model sophistication.

Large cloud providers need not choose only one architecture. Google offers both its own and NVIDIA accelerators and supports software portability between GPUs and TPUs. A growing market can support several routes. As work becomes mature and repetitive, customers have more scope to compare platform costs. The pressure on NVIDIA is competition for new orders and pricing, not AI suddenly losing its uses.

Making alternative hardware useful depends on adaptation and maintenance.

05

Customers pay for systems—and can build alternatives

NVIDIA’s development documentation describes two sets of tools: its Triton backend batches requests, reuses cached results and supports execution across devices; Dynamo plans capacity around request throughput and waiting-time targets. These tools aim to turn GPU computing capacity into a usable service.

What system functions can customers value?

Diagrama e comparação de dados
  1. 01Deploy changing models

    Reuse software, compilation and backends instead of adapting from scratch.

  2. 02Coordinate the system

    Batching, caching, nodes and networking work together.

  3. 03Meet service requirements

    Deliver specified quality, waiting times and operating conditions.

  4. 04Earn repeat purchases

    Orders, earnings and collection test customers’ willingness to keep paying.

The links combine NVIDIA development documentation and business analysis, not independent hardware tests or measured customer savings. Meta’s particular alternative configuration and AWS deployment illustrate buyer-led optimization.

KV caching saves intermediate results for reuse during later generation. Waiting-time targets cover how soon output starts and the intervals between subsequent outputs. The documentation describes functions, not measured customer savings. For customers repeatedly deploying new models, the tools’ value depends on how much adaptation they avoid and whether they meet service requirements.

Meta supplies a specific counterexample. It reported DDA communication optimization bringing MI300X to overall H100 performance parity in its particular test configuration. That does not settle every workload comparison. It shows an engineering-capable buyer improving an alternative instead of paying NVIDIA for all the integration. AWS’s in-house deployment is another route that depends on organizational capabilities.

NVIDIA must keep making its systems worth retaining when customers change models or add services. Faster deployment, less operating work and reliable completion can preserve orders. Customers will compare those benefits with system prices. A software ecosystem is a competitive asset, not a guarantee that profit remains safe.

Product documentation explains part of why customers might buy; the financial statements show how much NVIDIA earned this quarter.

06

Operating profit grows, but cash does not keep pace

NVIDIA’s operating income rose 19.0% sequentially in the quarter, but net income increased only 2.3%. Looking only at the bottom line can mistake nonoperating swings for a slower chip business. The difference largely comes from changes outside operating earnings.

Operating income increased $10.198 billion from the previous quarter; other income net fell $8.594 billion and tax expense rose $0.237 billion. Net income consequently increased $1.367 billion. The bridge below shows how lower nonoperating income held back net-income growth without establishing a stagnant chip business.

Lower other income restrains net-income growth

USD billions

FY2027 Q1 → Q2 · · Bilhões USD · trimestres fiscais individuais

Lower other income restrains net-income growthQuarterly NVIDIA net-income change in USD billions, not a cash bridge. +$10.198bn operating income − $8.594bn other income − $0.237bn tax expense = +$1.367bn net income. FMP MCP income records acquired October 9, 2026.58.321Prior netincome+10.198Operatingincrease-8.594Other incomedecline-0.237Tax expenseincrease59.688Current netincome
Ver os valores originais ↓
Lower other income restrains net-income growthUSD
Prior net income58,321,000,000
Operating increase10,198,000,000
Other income decline-8,594,000,000
Tax expense increase-237,000,000
Current net income59,688,000,000

Consultado em: 2026-10-09 · Evidências e guia de leitura ↗

Quarterly NVIDIA net-income change in USD billions, not a cash bridge. +$10.198bn operating income − $8.594bn other income − $0.237bn tax expense = +$1.367bn net income. FMP MCP income records acquired October 9, 2026.

Cash warrants separate attention. Q2 operating cash was $24.077 billion, down 52.2% sequentially but up 56.7% year over year. Receivables grew 54.9% sequentially, faster than revenue’s 17.9%. Recorded sales and profit increased without cash keeping pace. Collections, contract terms, company payment timing and working capital all matter to that difference.

Profit and cash moved differently this quarter, but the cash decline does not establish a shift to alternative chips. Turning profit into cash depends on customer collections, contract terms and NVIDIA’s own payments. One quarter’s cash decline does not establish shrinking inference demand.

07

Growing usage still leaves orders to compete for

Through the latest disclosed quarter, NVIDIA revenue and operating income were still growing, as was Google’s model API usage. These records do not show a simple outcome in which cheaper computation shrinks the business, nor do they establish lower costs as the cause of growth. For NVIDIA, the more useful question is whether that additional work will still lead customers to buy its systems.

Amazon’s operating Trainium deployment offers a concrete countercase. If large customers route more new inference work to in-house chips while NVIDIA’s gross profit and operating income remain under pressure, growing paid usage would not rescue the case for its system advantage. New orders continuing to choose NVIDIA alongside sustained operating-profit growth would strengthen that case. Whether customers pay within their contract terms requires a separate check.

Como esta página evoluirá
  • October 10, 2026: prose and figure explanations edited using the existing financial, usage and deployment sources.
  • Material changes in customer uses, hardware choices, quarterly earnings or collections will update the relevant explanations.
Evidências e guia de leitura

Financial figures cover NVIDIA’s consolidated business in standalone USD quarters. The main period ended July 26, 2026 and was disclosed August 26. FMP MCP acquisition dates: October 9 for income; October 8 for cash and balance records. Google and Amazon operating claims retain their service and time boundaries. Documentation and historical tests are not independently replicated customer results. The resource calculator uses assumptions. This business analysis does not forecast GPU units, industry elasticity or stock returns.

  1. FMP: NVIDIA quarterly income
    FMP: NVIDIA quarterly income ↗

    Company-wide revenue and earnings, rather than AI-segment profit or investment returns.

    Onde consultar · Quarterly income statements: revenue, gross profit, operating income, other income and tax expense.

    Consultado em: 2026-10-09 · Publicado: 2026-08-26
  2. FMP: NVIDIA quarterly cash flow
    FMP: NVIDIA quarterly cash flow ↗

    Company-wide US dollars, standalone quarters; both bridges use one consistent quarterly set.

    Onde consultar · Quarterly cash-flow statements: net income, noncash adjustments and working capital.

    Consultado em: 2026-10-08 · Publicado: 2026-08-26
  3. FMP: NVIDIA quarter-end receivables
    FMP: NVIDIA quarter-end receivables ↗

    Point-in-time receivables; growth does not measure actual collection days or delinquency.

    Onde consultar · Quarter-end balance sheets: accounts receivable.

    Consultado em: 2026-10-08 · Publicado: 2026-08-26
  4. Meta: scalingLLMinference
    Meta: scalingLLMinference ↗

    Specific buyer-reported MI300X/H100 test parity and 350/25ms timing goals. Not all workloads, TCO or achieved timing targets.

    Onde consultar · Inference communication bottlenecks; DDA performance tests on MI300X and H100.

    Consultado em: 2026-10-09 · Publicado: 2025-10-17
  5. NVIDIADynamoPlanner devdocs
    NVIDIADynamoPlanner devdocs ↗

    Dev documentation snapshot: throughput/SLA objectives and scale-down limitations. Not measured customer savings.

    Onde consultar · Capacity planning: throughput, first-token latency and output-token intervals.

    Consultado em: 2026-10-09
  6. NVIDIATritonTensorRT-LLM backend
    NVIDIATritonTensorRT-LLM backend ↗

    Backend capabilities and core-model benchmark excluding pre/post-processing latency. Not end-to-end business ROI.

    Onde consultar · TensorRT-LLM backend: batching, KV cache, multi-node execution and benchmark scope.

    Consultado em: 2026-10-09
  7. Google Q2 2026: CEO remarks on AI demand and efficiency
    Google Q2 2026: CEO remarks on AI demand and efficiency ↗

    Operator self-report: separate API usage, Search response costs and paid Cloud usage. Not a matched causal experiment or NVIDIA purchasing proxy.

    Onde consultar · Q2 2026 CEO remarks: model API usage, AI Mode, paid Cloud tokens and AI infrastructure.

    Consultado em: 2026-10-10 · Publicado: 2026-07-22
  8. Amazon: Trainium and production inference
    Amazon: Trainium and production inference ↗

    AWS operator statement, not independently reproduced share/TCO. Bedrock-specific majority, not global market share.

    Onde consultar · Production deployment: most Bedrock inference on Trainium; Claude training and operation on Trainium2.

    Consultado em: 2026-10-10 · Publicado: 2026-07-31
  9. OpenAI 2024: learning to reason with LLMs
    OpenAI 2024: learning to reason with LLMs ↗

    Historical research mechanism: more test-time computation can improve some reasoning tasks; not current model recommendation, demand elasticity or buyer ROI.

    Onde consultar · Learning to Reason with LLMs: performance changes with test-time computation.

    Consultado em: 2026-10-10 · Publicado: 2024-09-12

Redação DisconfirmAI · pesquisado e elaborado com IA. As fontes e o escopo dos cálculos estão listados acima.

Análises independentes aprofundadas

Estes estudos independentes examinam o fluxo de caixa, os recebimentos e as obrigações contratuais da NVIDIA. Escolha uma pergunta para ler a análise completa.

Estudo independente

NVIDIA profit barely changed. Why did operating cash fall by half?

In fiscal 2027 Q2, net income rose 2.3% from Q1 while operating cash fell 52.2%. A line-by-line reconciliation points to receivables, prepayments and payment timing. Extended terms for certain customers also complicate the idea that selling equipment means collecting cash first.

Ler a análise completa
Estudo independente

NVIDIA receivables outgrew sales. Does that establish customer trouble?

Receivables rose 54.9% sequentially against 17.9% revenue growth; five direct customers hold about 70% of the balance. These reveal capital absorption and concentration, not default.

Ler a análise completa
Estudo independente

What does NVIDIA carry under 90-day to one-year payment terms?

Certain large purchases separate delivery and payment. The first consequence is waiting for cash; trade credit, delinquency and conditional capacity purchases are separate matters.

Ler a análise completa
Estudo independente

Why can’t NVIDIA’s cash decline be assigned entirely to collections?

Prior-quarter liability support, prepayments and federal-payment timing provide competing explanations. Separate tax expense, payable and actual payments before assigning causes.

Ler a análise completa
Estudo independente

When could NVIDIA have to buy unsold AI-cloud capacity?

One type of agreement leaves unsold committed capacity for NVIDIA to purchase. Signing, triggering, execution and economic loss are separate states; aggregate commitments cannot collapse them.

Ler a análise completa
Estudo independente

After GPUs are sold, who still carries capital for AI compute?

The NVIDIA case shows collection waiting, conditional capacity buying and customer advances. These reveal contractual capital responsibility—not a verified industry cash loop.

Ler a análise completa
Temas de pesquisa →