Chinese open-weight models are cheap. Washington is deciding what that costs.
Moonshot AI released Kimi K3 on July 16 as the largest open-weight model yet released, reopening policy debates in Washington. Enterprises now question whether using Chinese open-weight models will remain feasible given potential regulatory changes.
NVIDIA Releases Cosmos 3 Edge: A 4B-Parameter Open World Model That Reasons and Generates Robot Actions On-Device
A 4-billion-parameter open world model called Cosmos 3 Edge enables robots and vision AI agents to understand surroundings and generate actions locally without cloud dependency. The Cosmos 3 family includes three versions: the 4B Edge model, a 16B Nano variant, and a 64B Super model released on May 31, 2026.
Atlassian: Research shows organizations should approach AI at the team level, not the individual level, to achieve true ROI
Eighty-nine percent of executives report individuals are speeding up with AI adoption, yet only 6% can identify specific ROI examples. Atlassian's research of 12,000 knowledge workers and 200 Fortune 1000 executives found that approximately 14% of teams translated AI usage into measurable value.
Writer's AI harness cuts token spend nearly 40% — without sacrificing accuracy
Token consumption for AI applications dropped nearly 40% through optimizations to the orchestration layer surrounding foundation models, without reducing accuracy. Engineering teams achieved these reductions by systematically refining the AI harness components that wrap the model, lowering cost-per-successful-task by up to 61%.
Alibaba’s Tongyi Lab Releases Qwen-Audio-3.0-TTS, a Hosted Text-to-Speech Model in Flash and Plus Tiers Across 16 Languages
Alibaba's Tongyi Lab released Qwen-Audio-3.0-TTS, a text-to-speech model available in Flash and Plus variants supporting 16 languages. The model is hosted through Alibaba Cloud Model Studio rather than offered as downloadable weights, with Flash optimized for real-time interaction and Plus for high-quality audio generation.
A single AI agent conversation can look perfect and still be broken, leaders from LangChain, Conviva and CoreWeave said at VB Transform 2026
Enterprises are shifting away from scoring individual AI agent conversations toward comparing cohorts of users against baselines, because single conversations can appear flawless yet indicate broken products. The industry faces tension between scalable but ungrounded automated judging through LLMs or agents and human review that cannot scale.
At VB Transform 2026, Zillow's engineering chief said AI ROI numbers only hold up if you measure before you build
Zillow built an AI architecture designed to maintain customer context across multiple touchpoints in real estate transactions rather than relying on isolated chatbots. The company discovered that establishing a persistent context layer proved more challenging than organizing its data infrastructure, despite initially focusing resources on data foundation and governance.
The cleanup trap: Stop asking RAG to fix bad data
Organizations mistakenly believe retrieval-augmented generation systems can fix fragmented legacy data at the retrieval layer, but production AI failures typically stem from inadequate enterprise data foundations rather than model limitations. The cleanup trap occurs when companies pipe ungoverned data into LLM orchestrators expecting the retrieval layer to patch problems through vector databases and embedding pipelines.
Safety guardrails blocked Hugging Face's defenders, not the attacker, when an AI agent breached its systems
Safety guardrails on frontier AI models blocked Hugging Face's incident response team from analyzing exploit data during a breach, while an autonomous AI agent moved undetected across the company's infrastructure for a weekend. Commercial models cannot cryptographically or organizationally distinguish between legitimate incident responders and malware authors requesting analysis.
The Download: AI hiring biases, and weather data sabotage
AI systems are more likely than humans to develop biases during hiring decisions. Résumés may be screened by artificial intelligence before human reviewers evaluate them.
Bristol Myers Squibb Building Life Science Industry’s Most Advanced AI Factory on NVIDIA Vera Rubin
Bristol Myers Squibb is deploying a second NVIDIA DGX SuperPOD to expand its AI computing capacity for drug discovery and development work. The company already operates one of the largest AI clusters in life sciences and is doubling its computational infrastructure investment.
US public health agencies to test OpenAI and Anthropic AI models
Ten state, local, tribal, or territorial jurisdictions will test generative AI models from OpenAI and Anthropic through a new program called PULSE. The initiative involves the Coalition for Health AI and Accenture to support public health departments in evaluating these AI tools.
Kimi K3 open-weight model: China’s biggest AI is a bet on memory, not compute
Moonshot AI released Kimi K3, an open-weight model with 2.8 trillion parameters, making it the largest open-weight model released to date. The model launched on July 16 and represents what the industry classifies as the 3T tier bracket.
AI is more likely than humans to form biases when hiring
Large language models develop biases when screening job applications, picking up prejudices from training data and creating additional ones independently. A study found AI systems show greater hiring discrimination than humans across multiple demographic categories in résumé evaluation.
AI confidence just dropped 17 points in six months. That’s actually great news.
IT leaders' assessment of their organizations' AI maturity fell from 40% to 23% over six months. Organizations deploying AI agents into production are revising downward because they encounter real-world operational challenges that pilots do not reveal.
Best Local LLMs You Can Run on a Single 24GB GPU in 2026: Qwen, Gemma, Mistral, DeepSeek Compared
Six open-weight models fit on a single 24GB GPU at Q4_K_M quantization, including Qwen3.6, Gemma 4, Mistral Small, gpt-oss-20b, and DeepSeek-R1-Distill. The guide evaluates each model's VRAM requirements, licensing terms, and performance strengths for local inference.
Feyn AI Releases SQRL, a Text-to-SQL Model Family That Inspects the Database Before Writing a Query
SQRL, a text-to-SQL model family, inspects databases with read-only probes before generating queries. The flagship SQRL-35B-A3B achieves 70.6% execution accuracy on BIRD Dev and distills into smaller 4B and 9B self-hostable versions.
Alibaba Previews Qwen3.8-Max, a 2.4 Trillion-Parameter Multimodal Model, Days After Moonshot’s Kimi K3 Open-Weight Launch
Alibaba previewed Qwen3.8-Max, a 2.4 trillion-parameter multimodal model available at 10% standard pricing on multiple platforms. The preview lacks published benchmark tables, model cards, licenses, per-token pricing, or active-parameter counts.
Perplexity AI Releases WANDR: An Open Benchmark Evaluating Research Agents That Must Search Wide And Deep
An open benchmark named WANDR containing 500 evidence-heavy tasks evaluates whether research agents can discover multiple qualifying entities with cited, re-verifiable evidence. Perplexity Search as Code achieved the highest scores at 0.363 soft F1 and 0.133 hard F1 on the evaluation harness.
Kimi K3 vs DeepSeek V4 Pro vs GLM-5.2: Open Trillion-Scale MoE Models Compared on Benchmarks, License, and Serving Cost
Three trillion-scale mixture-of-experts models with open weights were compared across benchmarks, licensing terms, and serving costs. The comparison evaluated Kimi K3, DeepSeek V4 Pro, and GLM-5.2 using measured intelligence scores and practical deployment expenses.
NVIDIA Released DeepStream 9.1: Bringing Agentic AI to Vision AI With 13 Skills and Multi-View 3D Tracking
DeepStream 9.1 introduces 13 agentic skills enabling coding agents to build multi-camera video analytics pipelines from natural-language prompts. Multi-View 3D Tracking fuses per-camera detections into a shared 3D world with globally consistent object IDs while AutoMagicCalib removes manual camera calibration.
Google Cloud’s Always-On Memory Agent Replaces RAG and Embeddings With Continuous LLM Consolidation on Gemini 3.1 Flash-Lite
An Always-On Memory Agent uses continuous LLM consolidation in SQLite rather than vector databases and embeddings to maintain persistent memory. Three sub-agents handle ingestion, consolidation, and querying of structured memory data running on Gemini 3.1 Flash-Lite.
Sakana AI’s Error Diffusion Trains Dale-Compliant Dual-Stream Networks, Reaching 96.7% MNIST and 61.7% CIFAR-10 Without Backpropagation
Error Diffusion trains dual-stream networks following Dale's principle without backpropagation, achieving 96.7% accuracy on MNIST and 61.7% on CIFAR-10. The method routes modulo errors through excitatory and inhibitory channels to scale learning across tasks including reinforcement learning.
Capital One releases VulnHunter, an open-source AI tool that finds software flaws before hackers do
An open-source AI tool called VulnHunter scans source code for exploitable vulnerabilities and proposes fixes before deployment. The tool uses "attacker-first forward analysis" to trace how adversaries would enter systems through APIs, network messages, or file uploads.
Intuit scrapped its own AI agent architecture twice in four months. At VB Transform 2026, its AI VP called that the fast path
Intuit rebuilt its AI agent architecture twice within four months, first consolidating specialist agents under a central orchestrator, then switching to a skills-based system after the orchestrator collapsed from complexity. The orchestrator failed because agents passing results in natural language lost critical context at each handoff, compounding errors across multiple agent interactions.
Agents think in milliseconds, legacy infrastructure doesn't. LinkedIn, Walmart and Zendesk shared how they closed the gap at VB Transform 2026
Enterprise infrastructure designed for human workflows became the bottleneck when scaling AI agents to production, not the underlying AI models themselves. Companies like LinkedIn, Walmart, and Zendesk each encountered different infrastructure failures when moving agents from pilot to production environments.
Brex built its AI agent policy by watching what agents actually do, not by writing rules first
Brex built CrabTrap, an open-source HTTP/HTTPS proxy that intercepts agent network traffic and uses a large language model to approve or deny requests based on policy rules. The platform monitors every API call agents make at the network layer rather than relying on traditional SDK-level permissions or model guardrails.
How Apple’s big lawsuit could disrupt OpenAI’s IPO plans
Apple filed a trade secrets lawsuit against OpenAI, alleging misconduct by its chief hardware officer and claiming over 400 former Apple employees work there. The lawsuit's timing potentially complicates OpenAI's reported plans to pursue an initial public offering.
Bunkerhill raises $55M to scale agentic AI across health systems
Bunkerhill Health raised $55 million in Series B funding to expand its Carebricks agentic AI platform across health systems. The round included participation from Sequoia Capital, Felicis, Optum Ventures, and Y Combinator.
Patreon stops asking AI bots not to scrape — and starts blocking them
Patreon partnered with Cloudflare to actively block AI bots from scraping creator content, shifting away from passive robots.txt files. The platform is now preventing unauthorized data collection for AI model training through active technical defenses.