Research
Notes for engineering and finance teams building AI spend governance.
Start here
Plain-English explainers if you're new to LLM cost.
- How do LLM providers charge?
- Input, output, cache, and reasoning tokens explained
- What is LLM cost attribution?
- Why does my LLM bill spike?
- How to budget for AI spend
All research
- /research/adaptive-thinking-cost-forecast.html
- A production runbook for an LLM cost spike
- A quarter of AI spend is slipping to 2027
- Add provenance to your AI output
- Agent Economics: What a Coding Agent Really Costs
- Agent spend attribution
- Agent spend guardrails
- Agent Spend Guardrails: Runtime Controls That Work
- Agent TCO: The 70% of Cost That Isn't Tokens
- AI Coding Economics 2026: GLM vs DeepSeek vs Claude
- AI Coding Plan Comparison
- AI Cost Allocation: Tags, Cost Centers and Chargebacks
- AI Cost Optimization Checklist for 2026
- AI Cost Optimization: Routing, Caching, Batching
- AI Cost Optimization: Six Levers That Actually Work
- AI cost recovery playbook: stop the bleeding
- AI FinOps: The Operating Model for GenAI Spend
- AI observability: cost, traffic, and quality · FinOps LLM
- AI provider list prices are not directly comparable
- AI spend variance: volume, rate, mix, and efficiency
- AI watermarking compliance cost
- Anthropic cost attribution
- Azure OpenAI vs Direct OpenAI Cost
- Batch APIs as a FinOps lever
- Caching strategies compared
- Chargeback vs Showback: Picking the Right FinOps Model
- Chargeback when your provider won't bill per team
- Cheapest Way to Run AI Code Generation (2026)
- Claude Fable 5.1 cost model
- Claude Max 20x vs Alibaba Token Pro Cost Audit
- Claude Sonnet 5 price rise cancelled: $2/$10 is permanent
- Cloud Infrastructure Cost Audit: Alibaba vs Claude Max
- Committed spend and reserved capacity discounts
- Cost per approved asset in image and video
- Cost per request as a product KPI
- Cost per successful task beats cost per request
- Currency exposure in LLM spend
- Cursor vs Copilot vs Claude Code Cost
- DeepSeek Harness Defensive Patterns
- DeepSeek Harness Goals and the Ralph Loop
- DeepSeek Harness Subagents and Delegation
- DeepSeek Harness Tool Scopes and Restrictions
- DeepSeek Harness: Everything Is a Plugin
- DeepSeek Harness: Install and First Run
- Eval cost allocation: who pays for LLM evals
- Finding the AI spend outside your API bill
- FinOps for LLM and GenAI: A Practical Framework
- FOCUS 1.5 adds AI token tracking
- Gartner's 2026 AI spending forecast
- GenAI cost management: controls and analytics (2026)
- GPT-5 Pricing and Budget Impact: A 2026 Planning Guide
- GPT-5 vs GPT-4o Cost Comparison: When to Upgrade
- GPT-5.6 Pricing Tier Guide: Luna, Terra, Sol
- Hidden LLM Costs Beyond Per-Token Pricing
- How AI watermarking works
- How AI watermarks get destroyed
- How do LLM providers charge? Tokens, tiers, caching
- How Much Does GPT-5 Cost? Complete 2026 Pricing Guide
- How prompting style shows up on the bill
- How to audit your LLM spend in 30 minutes
- How to budget for AI spend: a starter framework
- How to cap inference costs and prevent runaway spending
- How to prove AI ROI without hiding quality regressions
- Input, output, cache, reasoning: what tokens cost
- Invoice reconciliation for AI spend
- IT Chargeback and Showback Explained
- IT Chargeback and Showback for AI: A Practical Guide
- IT Showback: A Practical Guide for Technology Teams
- LLM API Pricing Tracker
- LLM Budget Governance: Alerts and Guardrails
- LLM Chargeback and Showback: Design Guide
- LLM cost anomaly detection
- LLM Cost Attribution by Team, Product and Tenant
- LLM Cost Benchmarks by Industry
- LLM Cost Calculator - Estimate Your Monthly AI Spend
- LLM cost dashboard: what to put on it
- LLM Cost Management: A Practical Framework
- LLM Cost Monitoring in Production: A Field Guide
- LLM cost monitoring: what to track and how to control it
- LLM Cost Per User Benchmarks
- LLM Cost Tracking: Tools, Tags and Telemetry
- LLM Cost Trends 2025–2026
- LLM Cost Visibility: Dashboards and Metrics That Matter
- LLM FinOps standards
- LLM Provider Arbitrage: Price and Quality Parity
- LLM Purchasing Guide for Platform Teams
- LLM Token Tracking by Input, Output and Cache
- LLM Usage Metering: How to Bill Internal AI Spend
- Long context vs RAG: where the cost breakeven actually is
- MCP Server Cost Impact: Tool Calls and Tokens
- Measure an AI feature's gross margin before you ship it
- Model deprecation migration: budgeting for forced moves
- Model Routing: Cheaper Models, Same Output Quality
- Multi-provider, multi-team LLM FinOps
- Multimodal cost allocation
- Non-production LLM spend
- On-prem and self-hosted LLM FinOps
- Open-Source vs Closed-Source LLM Cost Comparison
- OpenAI cost attribution
- OpenAI Fine-Tuning Sunset Economics
- OpenTelemetry GenAI conventions for cost attribution
- Ox Alpha was GLM-5.3
- Paying for the context window twice
- Pricing an open-weight model
- Prompt cache attribution
- Prompt caching explained
- Prompt Caching ROI: Calculating Real Savings
- Put cost controls inside your LLM evaluation pipeline
- RAG cost optimization
- Rate limits, 429s and tier upgrades
- Reasoning Model Cost Guide
- Reasoning token attribution and chargeback
- Regional and residency pricing premiums
- SaaS AI credits: metering a unit you do not control
- Semantic cache economics
- Server tools are a separate invoice
- Showback vs Chargeback: Which Model Fits Your Team?
- Spend Guardrails in DeepSeek Harness
- State of FinOps 2026: 98% now manage AI spend
- The DeepSeek Harness Session Log
- The first month-end close for an AI-heavy product
- The free tier as customer acquisition cost
- The real cost of switching LLM providers
- The recurring cost of embeddings and vector stores
- The retry loop that quietly adds to inference spend
- The stealth-model free-token playbook
- The three budgets every autonomous agent needs
- The token cost of JSON mode and schemas
- The Token Count Is Lying: Claude Max vs Qwen Audit
- The True Cost of Coding Agents: Beyond the API Bill
- The unit economics of realtime and audio APIs
- Token budget implementation
- Token budget implementation guide
- Token prices fell. Your LLM bill still went up.
- What guardrail and moderation passes cost
- What is AI value management?
- What Is DeepSeek Harness? A Complete Guide
- What is LLM cost attribution? A plain-English explainer
- What is LLM FinOps? A plain-English definition
- What is the Tokenomics Foundation?
- Why agentic AI blows the budget
- Why does my LLM bill spike? The most common causes
- Why Showback Comes Before Chargeback in FinOps
Want this applied to your own LLM spend? FinOps LLM runs a free audit of your AI costs and shows where the savings are. Book free audit →