Enterprise AI is entering an evaluation gap: Agents are gaining autonomy faster than companies can verify them
Half of surveyed enterprises deployed AI agents that passed internal evaluations but caused customer-facing failures, with 25% experiencing multiple failures. Despite this, 66% of companies permit or plan production deployment without human review, while only 5% trust their automated evaluations.
Read full story →Hyundai Motor Group Puts Data Flywheel Into Full Operation
Hyundai Motor Group activated its Data Flywheel system and demonstrated Level 2++ autonomous driving technology at an event in South Korea on September 13, 2026. The Group presented a dual-track autonomous driving roadmap alongside development strategies and implementation plans at 42dot headquarters in Gyeonggi Province.
Fly Language Model (FLM) Wires the Full Fruit Fly Connectome Into a Frozen 1.2B LLM, and Its Own Controls Show the Wiring Does Not Help
A fruit fly connectome with 166,700 neurons and 25.6 million connections was integrated into a 1.2 billion parameter frozen language model, training only 278,528 additional parameters. The connectome integration achieved 0.0222 nat per token improvement, but control models without the biological wiring performed slightly better across all random seeds.
Altman Says OpenAI Will Match Anthropic’s Embedded Evaluator Pledge
OpenAI committed to embedding independent evaluators with employee-level access to its systems, matching Anthropic's earlier pledge on the same day. Both companies endorsed slowing frontier AI development pace, with OpenAI promising to share additional details soon.