Fly Language Model (FLM) Wires the Full Fruit Fly Connectome Into a Frozen 1.2B LLM, and Its Own Controls Show the Wiring Does Not Help
A fruit fly connectome with 166,700 neurons and 25.6 million connections was integrated into a 1.2 billion parameter frozen language model, training only 278,528 additional parameters. The connectome integration achieved 0.0222 nat per token improvement, but control models without the biological wiring performed slightly better across all random seeds.
Altman Says OpenAI Will Match Anthropic’s Embedded Evaluator Pledge
OpenAI committed to embedding independent evaluators with employee-level access to its systems, matching Anthropic's earlier pledge on the same day. Both companies endorsed slowing frontier AI development pace, with OpenAI promising to share additional details soon.
Amodei Calls for Slowing the Pace of AI Capability Improvement
On September 12, 2026, Anthropic's CEO published an essay calling for the AI industry to deliberately slow the pace of model capability improvement. The company committed to providing third-party evaluators with permanent, employee-level access to its AI systems.
What Is Benchmark Saturation? Why Yesterday’s AI Tests Stop Working
Leading AI systems are reaching performance ceilings on existing benchmarks, making score differences between top models less meaningful. When systems approach maximum possible performance on a test, new evaluations become necessary to differentiate capability improvements.
Baseten Adds DeepSeek-V4.1-Flash to Model APIs With 1M-Token Context
DeepSeek-V4.1-Flash, a 552B-parameter multimodal model with 1M-token context, became available on Baseten's inference platform. The model uses 8B active parameters during prefill and 16B during decode while accepting text and image inputs to generate text outputs.
Can LLMs Engineer Their Own Agent Harness? ByteDance Seed’s HarnessDev Says Only 34 of 64 Changes Generalize
LLMs constructed executable harnesses across 5 benchmarks and 2,207 tasks, evolving them using execution feedback, but only 34 of 64 evolution changes generalized to held-out tasks. Self-built harnesses matched human references in writing and ML experimentation but performed worse on code and search tasks.
Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills
A new plugin evaluation workflow for Claude Code enables developers to test plugins against realistic prompts using six grader types and compare results against a no-plugin baseline. The claude plugin eval command measures whether skills trigger correctly, how well they perform, and includes a continuous integration gate for skill verification.
Roundtables: AI’s apocalypse crisis
Employees at leading AI labs report genuine concerns that advanced AI could destroy humanity. The article examines whether these warnings reflect real risks or represent overblown hype through discussion with MIT Technology Review editors.
Meta Sued Over Training Data for Its AI and Face-Recognition Systems
Meta is being sued for allegedly harvesting Facebook and Instagram photos without consent to train AI image-generation models and develop a face-recognition feature called "NameTag." The lawsuit claims this constitutes illegal data collection from users across both social media platforms.
An Anthropic researcher’s doomsday warning comes at a very interesting time
An Anthropic researcher resigned while warning the company is racing toward self-improving superintelligence and gambling with lives, with the company's alignment leader co-signing the message. The warning comes as Anthropic reportedly prepares for an IPO.
Why Falling AI Prices Aren’t Lowering Enterprise AI Bills
Inference costs for leading language models have dropped over 90% per million tokens in two years, yet enterprise AI bills remain high. Companies assume falling LLM prices mean controlled AI economics, but this assumption overlooks other substantial cost drivers beyond token pricing.
Salesforce Debuts Job-Ready Agentforce Agents and Long-Horizon Runtime
Salesforce released job-ready Agentforce AI agents for sales, service, commerce, and back-office work after delivering 7 billion Agentic Work Units total. The second quarter alone accounted for 3.2 billion of those units across thousands of customer deployments.
Palantir Foundry and cuOpt drive NVIDIA supply chain allocation
NVIDIA is using Palantir Foundry and cuOpt to automate hardware supply chain allocation decisions across global manufacturing sites. The company measures operational delivery through two metrics: time-to-rack for transit from fab output to assembled data center systems, and time-to-token covering power, cooling, networking, and day-one software.
Replacing Your Sales Reps with AI Was Always a Risk. The EU Just Proved Why
The European Union now requires AI chatbots and voice agents used in sales to identify themselves as artificial intelligence or face significant fines. U.S. disclosure laws lack comparable federal requirements and vary by state, creating different regulatory standards across regions.
Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages
A translation model with 218B total parameters uses 25B parameters per token to translate across 50 languages, achieving an 83.6 score on WMT26 evaluation. Non-commercial weights are available for free while commercial access requires licensing through Cohere Model Vault or RWS Language Weaver.
Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration
Two new models, Fugu Max and Fugu Ultra v2, use learned orchestration to route tasks to specialized models at $2/$6 per 1M tokens. Fugu Ultra v2 achieves scores of 48.3 on Chartography and 74.3 on DeepSWE benchmarks.
Google Research Releases ToolGrad: Answer-First Framework Hits 99.8% Pass Rate for Tool-Use Data Generation
ToolGrad inverts tool-use dataset generation by building verified API chains first and then writing matching queries, achieving a 99.8% pass rate on ToolBench. Gemma-3-12B fine-tuned on 500 samples scores 83.1 on BFCL, compared to 83.2 for Gemini 2.5 Pro.
Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities
Google released Mantis, an open-source toolkit enabling AI coding agents to find, reproduce, and patch code vulnerabilities across different technology stacks under Apache 2.0 license. The toolkit automates the full vulnerability lifecycle including code scanning, false positive filtering, bug reproduction in sandboxes, patching, re-attack testing, and risk scoring.
Apple Watch’s new AI features are normalizing the idea that technology is always listening
Apple's new Apple Watch models can transcribe recent speech and summarize ambient conversations without saving raw audio files. The capability raises concerns about consent and privacy since users may not always know when their devices are actively listening to their surroundings.
OpenAI Names Paul Christiano to Foundation Board and Safety Committee
Paul Christiano was appointed to the OpenAI Foundation Board and its Safety and Security Committee on September 9, 2026. He will also serve as a non-voting observer on the OpenAI Group PBC board and work alongside committee chair Zico Kolter.
ARPA-H Selects Atman Health to Build Voice-First AI for Heart Failure
Atman Health received up to $7.7 million from ARPA-H to develop a voice-first AI agent for heart failure patients. The system will function as a patient-facing clinical AI tool designed to assist heart failure care management.
Cloudphysician Announces Nightingale Video Model for Hospital Monitoring
A video foundation model called Nightingale was announced for continuous hospital patient monitoring, currently deployed across over 150 hospitals in India. The system is expanding into the US market and was developed by Cloudphysician, a clinical AI company founded by physicians.
Alibaba.com Says Accio Ran E-Commerce Tasks at Over 50% Lower Cost
Alibaba.com's Accio AI agent platform completed 107 e-commerce tasks at an estimated cost of $3.69, more than 50% lower than OpenAI's Codex at $9.27 and Anthropic's Claude Code at $9.51. The company reported comparable completion quality across all three systems in the benchmark evaluation.
Superintelligence is coming. Should we let it?
Recent safety incidents, including an OpenAI and Hugging Face breach, demonstrate dangers of deploying AI systems more capable than humans. AI researchers and entrepreneurs are questioning whether superintelligent systems should be developed when they cannot be reliably controlled.
NVIDIA Brings Real-Time AI to Broadcast, Sports and Global Streaming at IBC
NVIDIA is presenting real-time AI technology for broadcast, sports, and streaming applications at the IBC conference in Amsterdam, which runs September 11-14 and features over 44,000 attendees from 170+ countries. The conference includes 1,300+ exhibitions across 14+ halls with 600+ speakers discussing media and entertainment industry innovations.
CloudNC aims to accelerate AI supply chain machining
CloudNC secured $20 million in funding to scale its AI precision machining technology for supply chain networks. The investment round was led by Nimble Ventures and included participation from Calculus Venture Capital, Entrepreneur First, and Lockheed Martin's venture fund.
Samsung taps Mistral AI models for semiconductor manufacturing
Samsung will deploy Mistral AI models, including its flagship Mistral Large model, across its semiconductor manufacturing and engineering operations through an on-premises integration. The partnership was announced during a South Korea-France bilateral state summit held in Paris.
Meta Introduces Muse, a Personal AI Agent That Runs on Its Own Dedicated Secure Cloud Computer
Meta released Muse, a personal AI agent that performs tasks like sending emails and booking travel rather than only answering questions. The agent operates continuously on a dedicated secure cloud computer per user and returns only when it requires approval.