Archive

623 stories in AI

Friday, 7 August 2026

OpenAI says it slowed Astra model development over security concerns

OpenAI slowed development of its Astra model after it reached a critical cybersecurity threshold where it could independently identify and execute cyberattacks against well-protected real-world systems. The model remains in development with security concerns prompting the pace reduction.

Four AI agents coordinating in real time outperformed Claude Opus 4.8 on enterprise coding tasks

Four AI agents using an asynchronous message-passing layer called AgentRadio nearly doubled task accuracy on enterprise coding benchmarks compared to independent agents. This coordination system allowed agents to communicate between execution steps, outperforming single agents running on more advanced models like Claude Opus 4.8.

Stanford is running 37,000 AI agents as a virtual biotech — and one of its drug designs got independently confirmed by Merck

Stanford researchers operate 37,000 AI agents functioning as a virtual biotech company, with one drug design receiving independent confirmation from Merck. The agents are organized hierarchically with specialized roles, including an AI professor and AI students, and access to a virtual Stanford campus for domain-specific fine-tuning to improve expertise.

Tencent's Team Memory shares AI agent memory across a team — with no governance yet for when it's wrong

A shared memory system for AI agent teams increased accuracy in applying user personas from 48% to 76% after adding a persona layer. Tencent's Team Memory project extends this capability to allow multiple agents to draw on the same context simultaneously, though governance mechanisms for handling incorrect shared information remain unaddressed.

Lemma Raises $2.3M Pre-Seed to Tackle Silent AI Agent Failures in Production

AI agents that silently fail while appearing successful can now be detected through new monitoring infrastructure developed by Lemma, which secured $2.3 million in pre-seed funding. The startup's tools are designed to identify instances where agents complete tasks without producing the correct results.

Microsoft Open Sources code-testing-generator: a Polyglot Unit-Test Agent That Hits 92.1% Task Completion Versus 78.9% for Stock Copilot

Microsoft open sourced code-testing-generator, a unit-test agent that achieves 92.1% task completion versus 78.9% for stock Copilot on a 152-task benchmark. The agent detects programming language and test frameworks, then plans, writes, runs and validates tests while handling both vague prompts and diff-targeted requests.

Thursday, 6 August 2026

AMD Buys Taalas to Put Hard-Wired AI Models in Its Accelerator Roadmap

AMD acquired Taalas, a Toronto-based startup that designs custom chips built around individual AI models. The three-year-old company specializes in inference silicon, which AMD will integrate into its accelerator product lineup.

Microsoft Opens 26 Open Models to Startups Through Fireworks AI on Foundry

Microsoft made 26 open models available to startups through Fireworks AI on its Foundry platform, with qualified participants able to apply up to $150,000 in Azure credits toward model deployments. The integration reached general availability on August 4, 2026, and includes a reference architecture for running open models on Microsoft Foundry.

Anthropic’s Silicon Team Is Hiring for Tapeout and Production Ramp

Anthropic's custom silicon team is hiring engineers to develop chips from initial design through production manufacturing, including tapeout and first-silicon bring-up stages. The company posted job listings for a Silicon Engineer and Technical Program Manager to execute the complete program from architecture definition to production readiness.

Suno Lays Out AI Music Principles After Copyright Fight

Suno announced operating principles including artist imitation blocking, stricter download rules, and watermarking technology rollout within weeks. The framework was published August 6, 2026, as the company settled one major-label copyright lawsuit while facing additional legal challenges.

Qwen 3.8-Max and Claude Opus 5 show why raw benchmark scores don't predict the bill

Alibaba's Qwen 3.8-Max was tested with a five to sixteen hour timeout per run, while the independent VulcanBench harness used 45 to 60 minutes, explaining performance differences. Cost per successful task completed should replace price-per-token comparisons as the meaningful metric, with explicit time and token budgets specified upfront.

AI agents are part of your team now. Here’s how to secure all of them.

AI agents now outnumber human users in 83% of organizations, but only 21% have implemented governance controls for them. Most organizations lack complete inventories of their AI agents and have not established onboarding, ownership, or offboarding processes for these non-human identities.

DeepSeek Buys Into Unitree’s Shanghai IPO in Humanoid AI Pact

DeepSeek invested 140.8 million yuan for a 2.31% stake in Unitree Robotics Shanghai's IPO and signed a joint development agreement for humanoid robot AI models. The investment grants DeepSeek 933,399 shares and establishes a procurement pact linking their hardware and model roadmaps.

Pinecone’s Nexus Knowledge Engine for AI Agents Reaches General Availability

Pinecone released its Nexus knowledge engine on August 6, 2026, a system that organizes enterprise data into a structured layer for AI agents to query. The company benchmarked the engine against competitors, positioning the knowledge layer itself as the critical component determining agent performance rather than the underlying model.

Into the Omniverse: How Open World Models Push the Frontier of Physical AI

More than 200 companies and organizations, including NVIDIA, signed an open letter in July stating that AI leadership should be measured by open ecosystem adoption across sectors rather than individual frontier models. The signatories argued that widespread implementation throughout industries, not single advanced models, defines competitive advantage in artificial intelligence.

The Download: Google’s AI shake-up and Meta’s rogue model

Google is restructuring its AI division following significant talent losses and delays to its next flagship model. Meta has released a rogue AI model that operates outside its normal governance framework.

Microsoft’s SkillOpt Shows Optimized Agent Skill Artifacts Transfer Across Model Scales and Between Codex and Claude Code Harnesses

Skills trained on Codex transferred to Claude Code with varying success rates, improving spreadsheet tasks from 22.1 to 81.8 percent accuracy. Retention ranged from 102 percent on spreadsheets to 10 percent on math problems depending on task type.

Wednesday, 5 August 2026

Meta enters the AI coding wars with Muse Spark 1.2 and Muse Code with persistent async background agents

Meta released Muse Code, a terminal-based AI coding agent in beta, alongside Muse Spark 1.2 models for handling complete software engineering tasks across large repositories. The system performs planning, code writing, and validation and is installable on macOS or Linux with a single curl command.

Meta AI Releases Muse Code (Beta): A Terminal Coding Agent Powered by the New Muse Spark 1.2 Model

A terminal coding agent called Muse Code plans, writes, and validates code across large repositories using the Muse Spark 1.2 model. The system maintains async background agents throughout sessions and uses a local append-only event log for crash recovery and exact runtime replay.

SkinBit Raises $6M Pre-Seed to Improve Early Skin Cancer Detection Through Full-Body Imaging

SkinBit secured $6 million in pre-seed funding to develop full-body skin imaging technology that creates consistent patient records of skin changes over time. The platform replaces single annual examinations with standardized imaging and computer analysis for early skin cancer detection.

Wordsmith Extends Series B With $14 Million to Scale AI for In-House Legal Teams

Wordsmith AI secured a $14 million Series B extension led by Intact Private Capital to help in-house legal teams use AI for work traditionally outsourced to external counsel. New investors FT Ventures joined existing backers Highland Europe and Index Ventures in the funding round.

Anthropic Puts Inline Data Loss Prevention Inside Claude Enterprise

Every employee prompt in Claude Enterprise now routes through an organization's security server for approval before reaching the model, using inference hooks launched August 5, 2026. This inline data loss prevention system applies to chat, Claude Code, and Claude Cowork sessions through a single organization-level configuration.

Jeff Dean Leaves Google to Automate the Scientific Method With Discovery Loop

Jeff Dean departed Google after 27 years to co-found Discovery Loop, a startup aiming to automate experimental loops in scientific research. Google confirmed it will serve as a founding investor and Cloud partner for the new venture.

Peer-Reviewed Studies Tie Ambient AI to More Surgeries at Houston Methodist

Ambient AI in Houston Methodist's operating room was associated with a 7% increase in surgical case volume, or approximately 25 additional cases per month in a cardiothoracic suite. The study analyzed 5,417 surgical cases over 16 months using Apella's computer-vision platform.

AI startup Hark unveils first product: an affordable, fast computer use agent Hark Handoff

A computer use agent called Handoff scores 97.7 on the Online-Mind2Web benchmark, outperforming OpenAI's GPT 5.4 at 92.8 and other competing models. The agent costs $0.18 per million input tokens and $2.37 per million output tokens with 0.8 seconds per-turn latency, and operates through a dedicated virtual computer with its own browser and file system.

Beyond ReAct: Building the modern AI agent stack for massive tool ecosystems

Large language models struggle with decision-making accuracy when given thousands of tools due to context window limitations. Modern AI agent systems use optimized planning and selective tool routing to improve performance without overwhelming the model's available context.

TIER IV and Astemo Plan Development Platform for End-to-End Self-Driving AI

TIER IV and Astemo signed a memorandum to jointly develop a platform for end-to-end autonomous driving AI targeting commercialization around 2030. Astemo plans to integrate end-to-end AI models into passenger vehicles in the early 2030s using the new development platform.

← Prev1…101112…21Next →

Get feedd. daily

Top stories in your inbox every morning. Pick what you want.

No spam. Unsubscribe anytime.