Amazon AGI director says AI agent reliability, not capability, is blocking enterprise deployment at VB Transform 2026
Eighty-five percent of enterprises pilot AI agents, but only five percent deploy them to production because reliability issues cause agents to fail in real-world conditions despite passing internal evaluations. An Amazon AGI director identified four reliability dimensions,consistency, robustness, predictability, and safety,needed for enterprise deployment, citing an example where a QA agent read serial numbers correctly for two months before a software change made its vision encoder sensitive to screen position.
Read full story →Cerebras Runs OpenAI’s GPT-5.6 Sol at 750 Tokens Per Second in New Ultrafast Tier
Cerebras announced it is running OpenAI's GPT-5.6 Sol model at 750 output tokens per second on a new Ultrafast service tier, delivering up to 14 times faster processing than Standard. The Ultrafast tier launched as a limited preview in the OpenAI API for select customers on August 13, 2026.
Google’s Gemini 3.7 Flash targets coding and agents with a 50% introductory price cut
Google released Gemini 3.7 Flash with a focus on coding and agentic workflows at half price through end of 2026, with input tokens costing $0.75 per million and output tokens $3.75 per million. The model arrives three weeks after Gemini 3.6 Flash, driven by developer feedback and algorithmic improvements aimed at reducing retries and manual oversight.
DeepSeek Harness launches as open source rival to Claude Code, alongside V4-Pro on API with higher prices
DeepSeek launched DeepSeek-V4-Pro, an updated flagship model for agentic workloads, and DeepSeek Harness v0.1, an open-source agent harness available on GitHub under MIT license. API pricing shifts to peak and off-peak rates beginning August 16 at 16:00 UTC, with substantially higher prices than current rates.