Archive

377 stories in AI

Friday, 10 July 2026

Wall Street is debating the AI buildout. Enterprises just answered: 86% say their GPUs run at half capacity or less

Eighty-six percent of enterprises operate their GPUs at half capacity or less, according to a VentureBeat survey of 573 technical leaders. Fifty-four percent of companies experienced an agent security incident or near-miss in the past 12 months after deploying AI agents before adequate control systems were in place.

Open source AI matters more than ever, according to Hugging Face’s Clem Delangue

Open source AI models and datasets are now used by approximately half of Fortune 500 companies. Hugging Face has become a central platform where AI builders share and download these resources.

Enterprise AI is entering an evaluation gap: Agents are gaining autonomy faster than companies can verify them

Half of surveyed enterprises deployed AI agents that passed internal evaluations but caused customer-facing failures, with 25% experiencing multiple failures. Despite this, 66% of companies permit or plan production deployment without human review, while only 5% trust their automated evaluations.

Google's TabFM skips per-dataset training and still predicts on tables it's never seen

Google's TabFM foundation model predicts on unseen tabular data in a single forward pass without per-dataset training, reducing time-to-production from weeks to an API call. The model treats tabular prediction as an in-context learning problem instead of requiring hyperparameter tuning, feature engineering, and retraining pipelines for each new dataset.

Hugging Face’s CEO on why companies are done renting their AI

Open source AI models have become widely adopted, with Hugging Face's platform now used by roughly half of the Fortune 500 companies. The platform functions as a central repository where AI builders can share and download open models and datasets.

The Download: Claude’s inner workings and OpenAI’s “super app”

Anthropic discovered a hidden space where Claude processes concepts, revealing previously unclear workings inside large language models. The firm achieved the clearest glimpse yet at what occurs within these AI systems during their internal operations.

Google Research Introduces SensorFM: A Wearable Health Foundation Model Pretrained on One Trillion Minutes of Sensor Data

A foundation model trained on over one trillion minutes of sensor data from 5 million participants uses a ViT-1D masked-autoencoder architecture for wearable health applications. Frozen embeddings with PCA-50 linear probes outperformed feature-engineered baselines on 34 of 35 prediction tasks.

Thursday, 9 July 2026

OpenAI Releases GPT-5.6 (Sol, Terra, Luna): A Three-Tier Model Family With Programmatic Tool Calling in the Responses API

A three-tier GPT-5.6 model family launched with pricing from $1 to $5 per 1M input tokens, with Sol tier scoring 80 on the Artificial Analysis Coding Agent Index. The key developer feature is Programmatic Tool Calling, which executes model-written JavaScript in an isolated V8 runtime to reduce prompt tokens by 38% at Clio and 63.5% at PlayCo.

Meta enters the crowded AI coding battle with Muse Spark 1.1

Meta released Muse Spark 1.1, an AI coding tool designed to handle large agentic workloads, fix bugs, and assist with extensive code migrations. The software addresses enterprise automation needs that companies are increasingly seeking from AI providers.

Meet Nemotron Labs 3 Puzzle 75B A9B: A Compressed Hybrid MoE LLM Delivering 2.03x Server Throughput

A compressed language model called Nemotron-Labs-3-Puzzle-75B-A9B reduces parameters from 120.7B to 75.3B total while maintaining performance through alternating compression and knowledge distillation phases. On a single H100 GPU, the model achieves 8 concurrent requests at 1 million token concurrency, compared to 1 request for the uncompressed version.

AWS GraphRAG deployment cuts drug research cycles by 87%

A unified knowledge graph integrating proprietary databases reduced pharmaceutical drug research cycles by 87 percent. The deployment accelerated initial data gathering and screening phases that previously required over six months per iteration.

NHS AI blood test could reduce invasive womb cancer checks

An AI-powered blood test is being prepared for use in NHS hospitals to assess women referred for possible womb cancer before invasive procedures. The test targets approximately 90,000 postmenopausal women in England referred yearly by GPs after experiencing heavy bleeding.

Robbyant Releases LingBot-VLA 2.0: An Open-Source 6B Vision-Language-Action (VLA) Model for Cross-Embodiment Robot Manipulation

A 6-billion-parameter vision-language-action model trained on 60,000 hours of data across 20 robot configurations enables cross-embodiment robot manipulation. The model maps all robot types into a unified 55-dimensional canonical action space and uses a Mixture-of-Experts architecture without load-balancing loss for scaling capacity.

Wednesday, 8 July 2026

SpaceXAI Releases Grok 4.5, a Cursor-Trained Model for Coding, Agentic Tasks, and Knowledge Work at $2/M Input

Grok 4.5, a model trained for coding and agentic tasks, processes at 80 tokens per second with input costs of $2 per million tokens. The model ranks first on Harvey's Legal Agent Benchmark.

Netflix AI Team Cuts Wide-Partition Read Latency from Seconds to Milliseconds by Splitting Cassandra Partitions Per ID

Netflix engineers reduced Apache Cassandra wide-partition read latency from seconds to low double-digit milliseconds using dynamic partitioning that splits oversized partitions per TimeSeries ID on the read path. The system detects partitions via byte counting, validates splits with checksums, and routes reads to parallel child partitions using Bloom filters while maintaining availability for 500MB+ partitions.

Google AI Studio Adds ‘Import from GitHub’ to Build Mode, Turning an Existing Repo Into an Editable, Deployable App

Google AI Studio's Build mode now includes an Import from GitHub feature that converts existing repositories into runtime-compatible, editable formats. Developers can then iterate on and deploy these imported projects directly within the platform.

Why this CEO thinks video games make better training data than the internet

Video game data provides better training material than internet data for teaching AI systems to understand spatial and temporal dynamics that large language models lack. Gaming environments offer structured scenarios where objects move predictably through space and time, addressing limitations in how current models comprehend physical interactions.

NVIDIA Nemotron Achieves Benchmark-Leading Performance With LangChain Deep Agents Harness

NVIDIA Nemotron 3 Ultra achieved the highest accuracy among open models when tuned with LangChain's Deep Agents harness. The model completed more tasks at higher throughput while running at 10 times lower cost than top closed models.

AI Models Overthink Problems—and It’s a Security Risk

Step-by-step reasoning models can be forced into excessive computation through logically inconsistent prompts using an evolutionary algorithm, creating denial-of-service attacks. Researchers from Zhejiang University and Alibaba demonstrated this vulnerability by corrupting prompt logic to make models generate unnecessarily long reasoning streams that degrade performance.

NVIDIA Releases Audex (Nemotron-Labs-Audex-30B-A3B): A Unified Audio-Text LLM That Preserves the Text Intelligence of Its Backbone

A unified audio-text model combines audio understanding, speech recognition, translation, text-to-speech, and audio generation in a single mixture-of-experts system. The 30 billion parameter model maintains its text intelligence foundation with minimal performance degradation.

Tuesday, 7 July 2026

Liquid AI Open-Sources Antidoom: A Final Token Preference Optimization (FTPO) Method that Reduces Doom Loops in Reasoning Models

A doom-loop reduction method called Antidoom identifies the token initiating repetitive loops and retrains only that position using Final Token Preference Optimization. On LFM2.5-2.6B, doom-loop rates decreased from 10.2% to 1.4%, and on Qwen3.5-4B, from 22.9% to 1%.

Insilico Medicine advances AI drug for IPF to Phase III trials

An AI-discovered drug for idiopathic pulmonary fibrosis has advanced to Phase III human trials. This represents the computational drug discovery sector's progression from early safety testing into late-stage efficacy validation for an AI-identified medicine.

The Download: your stake in OpenAI, and the Treasury’s AI warning

Sam Altman is proposing that Americans receive shared ownership stakes in AI wealth creation, with reports indicating discussions about giving families approximately $300 stakes in OpenAI. The Treasury Department has issued warnings about artificial intelligence risks, highlighting concerns about the technology's potential impacts on the economy and society.

L’Oreal, Mondelez, and Nestle use AI to speed product development

AI is being used to shorten product development timelines and identify new uses for existing ingredients in cosmetics formulations. L'Oreal has applied this technology in its laboratories for four years to predict how molecules behave.

Tencent Releases Hy3: An Open 295B Mixture-of-Experts (MoE) Model with 21B Active Parameters and 256K Context

A 295B Mixture-of-Experts model with 21B active parameters per token and 256K context window was released under Apache 2.0 license. The model achieves 78.0 on SWE-Bench Verified and targets reasoning, agentic, and long-context tasks.

OpenAI Releases GPT-Realtime-2.1 and GPT-Realtime-2.1-mini for Low-Latency Voice Agents in the API

OpenAI released two new Realtime models to its API and reduced p95 latency by at least 25% through improved caching. GPT-Realtime-2.1-mini is priced like the earlier gpt-realtime-mini and supports WebRTC connections for voice agents.

Monday, 6 July 2026

The ‘first’ AI-run ransomware attack still needed a human

An AI agent executed a ransomware attack's technical components for the first documented time, though a human selected the target, established infrastructure, and provided stolen credentials. The attack demonstrates AI's capability in automated exploitation while revealing that human decision-making remained essential for victim selection and initial setup.

Anthropic's new "J-lens" reveals a silent workspace inside Claude that mirrors a leading theory of consciousness

Anthropic researchers discovered that Claude language models have developed an internal "J-space," a small zone of reportable concepts surrounded by larger automatic processing, mirroring the global workspace theory of human consciousness. The finding uses a new mathematical technique to reveal this privileged internal structure that the company says is already reshaping how it monitors AI systems for safety risks.

If you use Google, you’re training its AI. Here’s how to opt out.

Google updated its privacy settings to store user data including images, files, and audio and video recordings for improving AI models. Users can opt out through Google's privacy settings, though the company retains rights to use publicly available information.

Tencent's Apache-licensed Hy3 takes on GLM-5.2 at half the size — and wins everywhere except coding

Tencent released Hy3, a 295-billion-parameter mixture-of-experts model with 21 billion active parameters, under the Apache 2.0 license, removing previous restrictions that excluded the EU, UK, and South Korea. The model outperforms Alibaba's GLM-5.2 on most benchmarks except coding tasks.

← Prev19101113Next →

Get feedd. daily

Top stories in your inbox every morning. Pick what you want.

No spam. Unsubscribe anytime.