>samit_hota
Back to security news
SN-2026-259InformationalOpen

OpenAI Cuts GPT-5.6 Luna and Terra API Costs, Launches Sol Fast Mode

Samit Hota·
CVE ID
N/A
Affected Products / Orgs
OpenAI API, ChatGPT Work, Codex CLI, GPT-5.6 Luna, GPT-5.6 Terra, GPT-5.6 Sol
#news#vulnerability-disclosure#openai

OpenAI has significantly slashed API pricing for its GPT-5.6 Luna and Terra models while introducing a high-throughput API tier for GPT-5.6 Sol. The price adjustments, announced directly by OpenAI, lower operational costs for developers and enterprise security teams relying on OpenAI GPT-5.6 models for automated code analysis, vulnerability scanning pipelines, and autonomous agentic workflows.

Alongside the API price reductions, OpenAI updated its integration across developer tooling, transitioning automated code review capabilities in the ChatGPT desktop and web applications, as well as the Codex CLI, to the cheaper, higher-performing GPT-5.6 Luna model. These updates reduce the financial overhead of running continuous automated code audits while allowing organizations on ChatGPT Work and Codex subscriptions to execute substantially more workloads under their existing usage allowances.

Detailed Breakdown of GPT-5.6 Model Price Cuts

The most significant price reduction applies to GPT-5.6 Luna, OpenAI’s lightweight model variant. API pricing for Luna dropped by 80% across both input prompt processing and output token generation:

  • GPT-5.6 Luna Input Tokens: Decreased from $1.00 per million tokens down to $0.20 per million tokens.
  • GPT-5.6 Luna Output Tokens: Decreased from $6.00 per million tokens down to $1.20 per million tokens.

OpenAI’s mid-tier model, GPT-5.6 Terra, also received a 20% price reduction across its standard API endpoints:

  • GPT-5.6 Terra Input Tokens: Decreased from $2.50 per million tokens down to $2.00 per million tokens.
  • GPT-5.6 Terra Output Tokens: Decreased from $15.00 per million tokens down to $12.00 per million tokens.

According to OpenAI’s internal evaluation benchmarks, Luna ranks at the top of its internal intelligence index relative to compared models in its capability class, despite its lower cost per task execution. OpenAI attributed these efficiency gains across Luna and Terra to recent backend architectural improvements achieved while optimizing its high-end GPT-5.6 Sol model.

Introduction of GPT-5.6 Sol Fast Mode

To accommodate time-sensitive execution environments, OpenAI introduced a Fast mode API option for GPT-5.6 Sol. While standard pricing for Sol remains unchanged, the Fast mode option delivers processing speeds up to 2.5 times faster than standard API requests without sacrificing model reasoning or intelligence performance.

Sol Fast mode costs double the standard API rate for Sol. OpenAI noted that Fast mode is specifically tailored for latency-critical use cases—such as real-time interactive coding assistance, complex agentic workflows, and automated incident response research—where execution delay directly degrades operational effectiveness. For general-purpose, non-time-sensitive workloads, standard mode remains the recommended option.

Impact on Codex CLI, ChatGPT Work, and Security Automation

The restructured pricing directly affects enterprise software pipelines and automated review infrastructure. OpenAI confirmed that usage accounting for enterprise subscription tiers, including Codex and ChatGPT Work, has been adjusted to reflect the lower underlying token costs. As a result, automated tasks deduct fewer units from organization quotas, allowing security and development teams to accomplish greater volume under existing resource limits.

Key functional updates across developer tooling include:

  • Auto-review Engine Upgrade: Auto-review capabilities in both the ChatGPT application and Codex CLI have been upgraded from legacy GPT-5.4 models to GPT-5.6 Luna.
  • Cost Reduction for Code Audits: Switching the backend engine for Auto-review to GPT-5.6 Luna delivers an estimated ten-fold (10x) cost reduction for automated code review workloads.

From an application security perspective, cutting token overhead for lightweight models substantially changes the economics of AI-assisted code analysis. DevSecOps teams integrating LLM-based static analysis, pull request auditing, or automated remediation testing can run continuous checks across every commit without exceeding monthly API allocations.

Security Engineering Considerations for Enterprise Teams

While lower API pricing expands opportunities for LLM-driven security automation, engineering teams managing OpenAI API integrations should consider the following operational steps:

  • Audit Pipeline Configurations: Ensure automated security tooling and custom scripts using Codex CLI or API integrations are explicitly configured to leverage GPT-5.6 Luna rather than legacy GPT-5.4 endpoints to capture the 10x efficiency gain.
  • Evaluate Sol Fast Mode for SecOps Workflows: Reserve GPT-5.6 Sol Fast mode for latency-sensitive tasks such as active incident response triage, dynamic agentic routing, or real-time threat hunting where the 2.5x speed increase justifies the 2x API cost premium.
  • Update Token Monitoring Thresholds: Because enterprise quotas for ChatGPT Work and Codex now stretch further, adjust internal API usage alerts to maintain accurate visibility over token consumption patterns and prevent unexpected spikes from misconfigured automated loops.

Found something similar in your stack?

Let's find out before it becomes an incident.

Book an advisory call