CloudCodeTree LogoCloudCodeTree
AI NewsTutorialsAbout
CloudCodeTree Logo
CloudCodeTree
  • AI News
  • Tutorials
  • About
← Back to AI News
Inference Prices Keep Falling: OpenAI Slashes Luna 80% and Alibaba Debuts a 2.4T Rival

Photo: Steve A Johnson / Pexels

Inference Prices Keep Falling: OpenAI Slashes Luna 80% and Alibaba Debuts a 2.4T Rival

Chris Harper

1 min read

Aug 4, 2026 · 12:05 UTC

AI
News
LLM

Two moves in four days signal accelerating commoditization: OpenAI cuts its mid-tier GPT-5.6 Luna by 80%, and Alibaba drops a 2.4-trillion-parameter MoE model with top computer-use scores.

OpenAI (July 30): GPT-5.6 Luna dropped from $1/$6 to $0.20/$1.20 per million input/output tokens — an 80% cut. Terra fell 20% (now $2/$12). Sol stays at $5/$30. Luna is now competitive with budget-tier pricing, making it an interesting option for high-volume agentic pipelines where you want frontier reasoning at commodity cost.

Qwen3.8-Max (Aug 3): Alibaba released a 2.4-trillion-parameter MoE model (roughly 95B active parameters per token) under a permissive license. Alibaba's own benchmarks put it ahead of Fable 5 on OSWorld computer-use (86.1 vs 85.0) and ahead of Opus 4.8 on Terminal-Bench. SWE-bench and FrontierSWE still favor Fable 5 (Qwen scores 67.7/73.5 vs Fable 5's 80.0/88.8). Independent evals haven't landed yet — treat self-reported numbers with appropriate skepticism.

Why it matters: Luna's new price removes the cost argument against frontier models for agentic loops. Qwen3.8-Max is a credible open-weight alternative for computer-use agents — watch for independent benchmarks this week.

Sources: OpenAI cuts GPT-5.6 Luna 80% — VentureBeat · Qwen3.8-Max debut — VentureBeat · Qwen3.8-Max — MarkTechPost