Tag
#api-costs
Every story tagged api-costs, newest first.

AWS's new G7 instances won its own benchmark with half the GPUs
AWS benchmarked its Blackwell-based G7 instances against G5, G6 and G6e on 30B mixture-of-experts models. The two-GPU box beat the four-GPU boxes, the cheapest configuration and the fastest one turned out to be different machines, and the same model cost 2.5 times more per token on retrieval traffic than on chat.
Ahmad J · Sep 8, 2026 · 6 min read

Google's Lyria 3.5 has one price, and the 30 second clip stayed on Lyria 3
Lyria 3.5 is Google's new music model and it costs exactly what Lyria 3 Pro cost: $0.08 a song. The only 30 second rate, $0.04, still belongs to a model the same pricing page calls legacy.
Ahmad J · Sep 6, 2026 · 6 min read

OpenAI's cyber model has one price and no way to lower it
GPT-5.6 Cyber is absent from all four of OpenAI's rate tables. It has a single set of rates, no batch discount, and you have to be approved before you can pay them.
Ahmad J · Sep 5, 2026 · 8 min read

GPT-6 Astra has a million-token window and a price cliff at 272,000
OpenAI's newest flagship advertises a 1,050,000-token context window. Cross 272,000 input tokens by one token and the entire request reprices at 2x input and 1.5x output. The worked example, the multiplier stack, and the four things that actually cut the bill.
Ahmad J · Sep 5, 2026 · 6 min read

Gemini Flash is half price until 31 December, and Google published the date it ends
Google's pricing page lists 32 Gemini models. Only three carry a second, dated price, and all three are the current Flash line, doubling on 1 January 2027. The comparison that makes it fair, the model card warning that compounds it, and what to budget instead.
Ahmad J · Sep 3, 2026 · 6 min read
