Stay ahead of the curve with the latest from artificial intelligence. We track model releases, research breakthroughs, industry moves, and real-world deployments across OpenAI, Google DeepMind, Anthropic, Meta AI, and the broader ML ecosystem, summarised in seconds.
Hyundai Motor Group Puts Data Flywheel Into Full Operation
Hyundai Motor Group activated its Data Flywheel system and demonstrated Level 2++ autonomous driving technology at an event in South Korea on September 13, 2026. The Group presented a dual-track autonomous driving roadmap alongside development strategies and implementation plans at 42dot headquarters in Gyeonggi Province.
Fly Language Model (FLM) Wires the Full Fruit Fly Connectome Into a Frozen 1.2B LLM, and Its Own Controls Show the Wiring Does Not Help
A fruit fly connectome with 166,700 neurons and 25.6 million connections was integrated into a 1.2 billion parameter frozen language model, training only 278,528 additional parameters. The connectome integration achieved 0.0222 nat per token improvement, but control models without the biological wiring performed slightly better across all random seeds.
Altman Says OpenAI Will Match Anthropic’s Embedded Evaluator Pledge
OpenAI committed to embedding independent evaluators with employee-level access to its systems, matching Anthropic's earlier pledge on the same day. Both companies endorsed slowing frontier AI development pace, with OpenAI promising to share additional details soon.
Amodei Calls for Slowing the Pace of AI Capability Improvement
On September 12, 2026, Anthropic's CEO published an essay calling for the AI industry to deliberately slow the pace of model capability improvement. The company committed to providing third-party evaluators with permanent, employee-level access to its AI systems.
What Is Benchmark Saturation? Why Yesterday’s AI Tests Stop Working
Leading AI systems are reaching performance ceilings on existing benchmarks, making score differences between top models less meaningful. When systems approach maximum possible performance on a test, new evaluations become necessary to differentiate capability improvements.
Baseten Adds DeepSeek-V4.1-Flash to Model APIs With 1M-Token Context
DeepSeek-V4.1-Flash, a 552B-parameter multimodal model with 1M-token context, became available on Baseten's inference platform. The model uses 8B active parameters during prefill and 16B during decode while accepting text and image inputs to generate text outputs.
Can LLMs Engineer Their Own Agent Harness? ByteDance Seed’s HarnessDev Says Only 34 of 64 Changes Generalize
LLMs constructed executable harnesses across 5 benchmarks and 2,207 tasks, evolving them using execution feedback, but only 34 of 64 evolution changes generalized to held-out tasks. Self-built harnesses matched human references in writing and ML experimentation but performed worse on code and search tasks.
Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills
A new plugin evaluation workflow for Claude Code enables developers to test plugins against realistic prompts using six grader types and compare results against a no-plugin baseline. The claude plugin eval command measures whether skills trigger correctly, how well they perform, and includes a continuous integration gate for skill verification.
Roundtables: AI’s apocalypse crisis
Employees at leading AI labs report genuine concerns that advanced AI could destroy humanity. The article examines whether these warnings reflect real risks or represent overblown hype through discussion with MIT Technology Review editors.
Meta Sued Over Training Data for Its AI and Face-Recognition Systems
Meta is being sued for allegedly harvesting Facebook and Instagram photos without consent to train AI image-generation models and develop a face-recognition feature called "NameTag." The lawsuit claims this constitutes illegal data collection from users across both social media platforms.