feedd.AI
AIVentureBeat · 17h ago

DeepSeek's top-ranked V4 Flash stumbles on real agent tasks as its prices surge

DeepSeek's V4 Flash completed only 53.8% of complex agent tasks across eight different harnesses testing 30 multi-step workflows involving live tools like Gmail and GitHub. The model produced substantially different results depending on the harness, tool configuration, caching behavior, and provider stack used.

Read full story →
More from AI

An evaluation harness discovered that AI models express highest confidence precisely when their answers are wrong, a problem qualitative review missed. This gap between outputs that sound correct and outputs that are verifiably accurate becomes critical as LLM tools influence business decisions in compliance, data analysis, and operations.

01

Chinese AI startup Z.ai released GLM-5.3, which demonstrates improved long-horizon coding and advanced cybersecurity capabilities through scaled post-training rather than new pretraining. The model reportedly identified a potentially serious vulnerability in Cursor, an AI coding startup acquired by SpaceX, with API access and open weights coming approximately two weeks after launch once safety evaluation completes.

02

Enterprise revenue has surpassed consumer revenue at OpenAI, occurring earlier than the company's previous forecast for late 2026. Finance chief Sarah Friar disclosed this crossover at an August 14, 2026 investor meeting, with the enterprise operation now generating more income than the ChatGPT consumer business.

03

Get feedd. daily

Top stories in your inbox every morning. Pick what you want.

No spam. Unsubscribe anytime.