DeepSeek cut prices 75%. The 100x problem remains
DeepSeek reduced V4-Pro pricing by 75 percent, but agent systems consume tokens at rates faster than prices decline, creating substantial cost pressures despite cheaper models. A single user request in agentic workflows generates multiple billable operations through planning, retrieval, tool use, verification, and summarization loops, multiplying costs by approximately 100 times compared to simple chatbot responses.
Mira Murati’s Thinking Machines Lab Makes The Technical Case For Human-Centered AI Built On Customizable Model Weights
Thinking Machines Lab published an essay proposing human-centered AI through decentralized model ownership and alignment as technical challenges. The framework uses LoRA fine-tuning methods that allow teams to train and retain their own customizable model weights.
A Coding Guide to NVIDIA’s Tile-Based GPU Programming: From cuTile and Triton Kernels to Flash Attention
NVIDIA's tile-based GPU programming uses TileGym to operate on whole data tiles rather than single threads, with implementations falling back from cuTile to Triton kernels depending on available hardware. The tutorial covers vector addition, fused GELU, row-wise softmax, tiled matrix multiplication, and flash attention operations validated against PyTorch.
57% of enterprises have watched AI agents be confidently wrong. The fix is an agentic context layer, but who has one?
Fifty-seven percent of enterprises reported AI agents providing confidently incorrect answers due to missing or inconsistent business context in the past six months. Thirty-eight percent of enterprises rely on document retrieval as their primary method for giving agents business context, while seventy-five percent lack a dedicated agentic context layer to govern this information.
OpenAI introduces ChatGPT Work, a cloud-based AI agent that manages tasks across email, Slack and calendars
OpenAI launched ChatGPT Work, an AI agent powered by GPT-5.6 that executes complex multi-step tasks across email, Slack, calendars, and code repositories. The agent can gather context from connected apps to produce finished documents, spreadsheets, presentations, reports, and websites while working independently on projects for hours.
Wall Street is debating the AI buildout. Enterprises just answered: 86% say their GPUs run at half capacity or less
Eighty-six percent of enterprises operate their GPUs at half capacity or less, according to a VentureBeat survey of 573 technical leaders. Fifty-four percent of companies experienced an agent security incident or near-miss in the past 12 months after deploying AI agents before adequate control systems were in place.
Open source AI matters more than ever, according to Hugging Face’s Clem Delangue
Open source AI models and datasets are now used by approximately half of Fortune 500 companies. Hugging Face has become a central platform where AI builders share and download these resources.
Enterprise AI is entering an evaluation gap: Agents are gaining autonomy faster than companies can verify them
Half of surveyed enterprises deployed AI agents that passed internal evaluations but caused customer-facing failures, with 25% experiencing multiple failures. Despite this, 66% of companies permit or plan production deployment without human review, while only 5% trust their automated evaluations.
Google's TabFM skips per-dataset training and still predicts on tables it's never seen
Google's TabFM foundation model predicts on unseen tabular data in a single forward pass without per-dataset training, reducing time-to-production from weeks to an API call. The model treats tabular prediction as an in-context learning problem instead of requiring hyperparameter tuning, feature engineering, and retraining pipelines for each new dataset.
Hugging Face’s CEO on why companies are done renting their AI
Open source AI models have become widely adopted, with Hugging Face's platform now used by roughly half of the Fortune 500 companies. The platform functions as a central repository where AI builders can share and download open models and datasets.
The Download: Claude’s inner workings and OpenAI’s “super app”
Anthropic discovered a hidden space where Claude processes concepts, revealing previously unclear workings inside large language models. The firm achieved the clearest glimpse yet at what occurs within these AI systems during their internal operations.
Google Research Introduces SensorFM: A Wearable Health Foundation Model Pretrained on One Trillion Minutes of Sensor Data
A foundation model trained on over one trillion minutes of sensor data from 5 million participants uses a ViT-1D masked-autoencoder architecture for wearable health applications. Frozen embeddings with PCA-50 linear probes outperformed feature-engineered baselines on 34 of 35 prediction tasks.
OpenAI Releases GPT-5.6 (Sol, Terra, Luna): A Three-Tier Model Family With Programmatic Tool Calling in the Responses API
A three-tier GPT-5.6 model family launched with pricing from $1 to $5 per 1M input tokens, with Sol tier scoring 80 on the Artificial Analysis Coding Agent Index. The key developer feature is Programmatic Tool Calling, which executes model-written JavaScript in an isolated V8 runtime to reduce prompt tokens by 38% at Clio and 63.5% at PlayCo.
Meta enters the crowded AI coding battle with Muse Spark 1.1
Meta released Muse Spark 1.1, an AI coding tool designed to handle large agentic workloads, fix bugs, and assist with extensive code migrations. The software addresses enterprise automation needs that companies are increasingly seeking from AI providers.
Meet Nemotron Labs 3 Puzzle 75B A9B: A Compressed Hybrid MoE LLM Delivering 2.03x Server Throughput
A compressed language model called Nemotron-Labs-3-Puzzle-75B-A9B reduces parameters from 120.7B to 75.3B total while maintaining performance through alternating compression and knowledge distillation phases. On a single H100 GPU, the model achieves 8 concurrent requests at 1 million token concurrency, compared to 1 request for the uncompressed version.
AWS GraphRAG deployment cuts drug research cycles by 87%
A unified knowledge graph integrating proprietary databases reduced pharmaceutical drug research cycles by 87 percent. The deployment accelerated initial data gathering and screening phases that previously required over six months per iteration.
NHS AI blood test could reduce invasive womb cancer checks
An AI-powered blood test is being prepared for use in NHS hospitals to assess women referred for possible womb cancer before invasive procedures. The test targets approximately 90,000 postmenopausal women in England referred yearly by GPs after experiencing heavy bleeding.
Robbyant Releases LingBot-VLA 2.0: An Open-Source 6B Vision-Language-Action (VLA) Model for Cross-Embodiment Robot Manipulation
A 6-billion-parameter vision-language-action model trained on 60,000 hours of data across 20 robot configurations enables cross-embodiment robot manipulation. The model maps all robot types into a unified 55-dimensional canonical action space and uses a Mixture-of-Experts architecture without load-balancing loss for scaling capacity.
SpaceXAI Releases Grok 4.5, a Cursor-Trained Model for Coding, Agentic Tasks, and Knowledge Work at $2/M Input
Grok 4.5, a model trained for coding and agentic tasks, processes at 80 tokens per second with input costs of $2 per million tokens. The model ranks first on Harvey's Legal Agent Benchmark.
Netflix AI Team Cuts Wide-Partition Read Latency from Seconds to Milliseconds by Splitting Cassandra Partitions Per ID
Netflix engineers reduced Apache Cassandra wide-partition read latency from seconds to low double-digit milliseconds using dynamic partitioning that splits oversized partitions per TimeSeries ID on the read path. The system detects partitions via byte counting, validates splits with checksums, and routes reads to parallel child partitions using Bloom filters while maintaining availability for 500MB+ partitions.
Google AI Studio Adds ‘Import from GitHub’ to Build Mode, Turning an Existing Repo Into an Editable, Deployable App
Google AI Studio's Build mode now includes an Import from GitHub feature that converts existing repositories into runtime-compatible, editable formats. Developers can then iterate on and deploy these imported projects directly within the platform.
Why this CEO thinks video games make better training data than the internet
Video game data provides better training material than internet data for teaching AI systems to understand spatial and temporal dynamics that large language models lack. Gaming environments offer structured scenarios where objects move predictably through space and time, addressing limitations in how current models comprehend physical interactions.
NVIDIA Nemotron Achieves Benchmark-Leading Performance With LangChain Deep Agents Harness
NVIDIA Nemotron 3 Ultra achieved the highest accuracy among open models when tuned with LangChain's Deep Agents harness. The model completed more tasks at higher throughput while running at 10 times lower cost than top closed models.
AI Models Overthink Problems—and It’s a Security Risk
Step-by-step reasoning models can be forced into excessive computation through logically inconsistent prompts using an evolutionary algorithm, creating denial-of-service attacks. Researchers from Zhejiang University and Alibaba demonstrated this vulnerability by corrupting prompt logic to make models generate unnecessarily long reasoning streams that degrade performance.
NVIDIA Releases Audex (Nemotron-Labs-Audex-30B-A3B): A Unified Audio-Text LLM That Preserves the Text Intelligence of Its Backbone
A unified audio-text model combines audio understanding, speech recognition, translation, text-to-speech, and audio generation in a single mixture-of-experts system. The 30 billion parameter model maintains its text intelligence foundation with minimal performance degradation.
Liquid AI Open-Sources Antidoom: A Final Token Preference Optimization (FTPO) Method that Reduces Doom Loops in Reasoning Models
A doom-loop reduction method called Antidoom identifies the token initiating repetitive loops and retrains only that position using Final Token Preference Optimization. On LFM2.5-2.6B, doom-loop rates decreased from 10.2% to 1.4%, and on Qwen3.5-4B, from 22.9% to 1%.
Insilico Medicine advances AI drug for IPF to Phase III trials
An AI-discovered drug for idiopathic pulmonary fibrosis has advanced to Phase III human trials. This represents the computational drug discovery sector's progression from early safety testing into late-stage efficacy validation for an AI-identified medicine.
The Download: your stake in OpenAI, and the Treasury’s AI warning
Sam Altman is proposing that Americans receive shared ownership stakes in AI wealth creation, with reports indicating discussions about giving families approximately $300 stakes in OpenAI. The Treasury Department has issued warnings about artificial intelligence risks, highlighting concerns about the technology's potential impacts on the economy and society.
L’Oreal, Mondelez, and Nestle use AI to speed product development
AI is being used to shorten product development timelines and identify new uses for existing ingredients in cosmetics formulations. L'Oreal has applied this technology in its laboratories for four years to predict how molecules behave.