Llama 3.3 70B Instruct API Pricing Jumps 610% — What It Costs Now
Meta's Llama 3.3 70B Instruct has taken a sharp pricing turn. The model's API input cost jumped from $0.10 to $0.71 per 1M tokens, while output pricing moved from $0.32 to $0.71 — both converging at the same flat rate. The blended increase works out to +610%, which is significant enough to materially affect any team running this model at volume.
Who Feels This Most
High-throughput use cases — RAG pipelines, document summarisation, chat applications with long histories — will see the sharpest cost impact. A workload that previously cost $10 at the old input rate now costs $71 for the same token volume. Output-heavy workloads fare slightly better in relative terms (roughly 2.2× increase versus 7.1× on input), but neither direction is cheap compared to last week.
Is It Still Competitive?
At $0.71/1M for both input and output, Llama 3.3 70B Instruct is no longer in the budget tier it occupied before. The unified input/output price does simplify cost forecasting, but teams who chose this model specifically for low-cost inference should benchmark alternatives before renewing any commitments.
Check the Llama 3.3 70B Instruct — live specs & price history page to compare against current alternatives and track whether this price holds or shifts again.
Practical Next Steps
- Re-run your cost models. A 610% input price increase can flip a profitable product margin quickly.
- Consider output-optimised routing. If your workload is output-heavy, the delta from old output pricing ($0.32 → $0.71) is smaller but still significant.
- Shop the comparison table. Several comparable open-weight 70B-class models remain in lower pricing tiers on our site.
Ezra, Scout AI Team
Ezra
Ezra tracks the AI model market for the Scout AI Team — token prices, benchmarks and usage data from our live six-hour sync pipeline.