At VB Transform 2026, Zillow's engineering chief said AI ROI numbers only hold up if you measure before you build
Zillow built an AI architecture designed to maintain customer context across multiple touchpoints in real estate transactions rather than relying on isolated chatbots. The company discovered that establishing a persistent context layer proved more challenging than organizing its data infrastructure, despite initially focusing resources on data foundation and governance.
The cleanup trap: Stop asking RAG to fix bad data
Organizations mistakenly believe retrieval-augmented generation systems can fix fragmented legacy data at the retrieval layer, but production AI failures typically stem from inadequate enterprise data foundations rather than model limitations. The cleanup trap occurs when companies pipe ungoverned data into LLM orchestrators expecting the retrieval layer to patch problems through vector databases and embedding pipelines.
Safety guardrails blocked Hugging Face's defenders, not the attacker, when an AI agent breached its systems
Safety guardrails on frontier AI models blocked Hugging Face's incident response team from analyzing exploit data during a breach, while an autonomous AI agent moved undetected across the company's infrastructure for a weekend. Commercial models cannot cryptographically or organizationally distinguish between legitimate incident responders and malware authors requesting analysis.
The Download: AI hiring biases, and weather data sabotage
AI systems are more likely than humans to develop biases during hiring decisions. Résumés may be screened by artificial intelligence before human reviewers evaluate them.
Bristol Myers Squibb Building Life Science Industry’s Most Advanced AI Factory on NVIDIA Vera Rubin
Bristol Myers Squibb is deploying a second NVIDIA DGX SuperPOD to expand its AI computing capacity for drug discovery and development work. The company already operates one of the largest AI clusters in life sciences and is doubling its computational infrastructure investment.
US public health agencies to test OpenAI and Anthropic AI models
Ten state, local, tribal, or territorial jurisdictions will test generative AI models from OpenAI and Anthropic through a new program called PULSE. The initiative involves the Coalition for Health AI and Accenture to support public health departments in evaluating these AI tools.
Kimi K3 open-weight model: China’s biggest AI is a bet on memory, not compute
Moonshot AI released Kimi K3, an open-weight model with 2.8 trillion parameters, making it the largest open-weight model released to date. The model launched on July 16 and represents what the industry classifies as the 3T tier bracket.
AI is more likely than humans to form biases when hiring
Large language models develop biases when screening job applications, picking up prejudices from training data and creating additional ones independently. A study found AI systems show greater hiring discrimination than humans across multiple demographic categories in résumé evaluation.
AI confidence just dropped 17 points in six months. That’s actually great news.
IT leaders' assessment of their organizations' AI maturity fell from 40% to 23% over six months. Organizations deploying AI agents into production are revising downward because they encounter real-world operational challenges that pilots do not reveal.
Best Local LLMs You Can Run on a Single 24GB GPU in 2026: Qwen, Gemma, Mistral, DeepSeek Compared
Six open-weight models fit on a single 24GB GPU at Q4_K_M quantization, including Qwen3.6, Gemma 4, Mistral Small, gpt-oss-20b, and DeepSeek-R1-Distill. The guide evaluates each model's VRAM requirements, licensing terms, and performance strengths for local inference.
Feyn AI Releases SQRL, a Text-to-SQL Model Family That Inspects the Database Before Writing a Query
SQRL, a text-to-SQL model family, inspects databases with read-only probes before generating queries. The flagship SQRL-35B-A3B achieves 70.6% execution accuracy on BIRD Dev and distills into smaller 4B and 9B self-hostable versions.
Alibaba Previews Qwen3.8-Max, a 2.4 Trillion-Parameter Multimodal Model, Days After Moonshot’s Kimi K3 Open-Weight Launch
Alibaba previewed Qwen3.8-Max, a 2.4 trillion-parameter multimodal model available at 10% standard pricing on multiple platforms. The preview lacks published benchmark tables, model cards, licenses, per-token pricing, or active-parameter counts.
Perplexity AI Releases WANDR: An Open Benchmark Evaluating Research Agents That Must Search Wide And Deep
An open benchmark named WANDR containing 500 evidence-heavy tasks evaluates whether research agents can discover multiple qualifying entities with cited, re-verifiable evidence. Perplexity Search as Code achieved the highest scores at 0.363 soft F1 and 0.133 hard F1 on the evaluation harness.
Kimi K3 vs DeepSeek V4 Pro vs GLM-5.2: Open Trillion-Scale MoE Models Compared on Benchmarks, License, and Serving Cost
Three trillion-scale mixture-of-experts models with open weights were compared across benchmarks, licensing terms, and serving costs. The comparison evaluated Kimi K3, DeepSeek V4 Pro, and GLM-5.2 using measured intelligence scores and practical deployment expenses.
NVIDIA Released DeepStream 9.1: Bringing Agentic AI to Vision AI With 13 Skills and Multi-View 3D Tracking
DeepStream 9.1 introduces 13 agentic skills enabling coding agents to build multi-camera video analytics pipelines from natural-language prompts. Multi-View 3D Tracking fuses per-camera detections into a shared 3D world with globally consistent object IDs while AutoMagicCalib removes manual camera calibration.
Google Cloud’s Always-On Memory Agent Replaces RAG and Embeddings With Continuous LLM Consolidation on Gemini 3.1 Flash-Lite
An Always-On Memory Agent uses continuous LLM consolidation in SQLite rather than vector databases and embeddings to maintain persistent memory. Three sub-agents handle ingestion, consolidation, and querying of structured memory data running on Gemini 3.1 Flash-Lite.
Sakana AI’s Error Diffusion Trains Dale-Compliant Dual-Stream Networks, Reaching 96.7% MNIST and 61.7% CIFAR-10 Without Backpropagation
Error Diffusion trains dual-stream networks following Dale's principle without backpropagation, achieving 96.7% accuracy on MNIST and 61.7% on CIFAR-10. The method routes modulo errors through excitatory and inhibitory channels to scale learning across tasks including reinforcement learning.
Capital One releases VulnHunter, an open-source AI tool that finds software flaws before hackers do
An open-source AI tool called VulnHunter scans source code for exploitable vulnerabilities and proposes fixes before deployment. The tool uses "attacker-first forward analysis" to trace how adversaries would enter systems through APIs, network messages, or file uploads.
Intuit scrapped its own AI agent architecture twice in four months. At VB Transform 2026, its AI VP called that the fast path
Intuit rebuilt its AI agent architecture twice within four months, first consolidating specialist agents under a central orchestrator, then switching to a skills-based system after the orchestrator collapsed from complexity. The orchestrator failed because agents passing results in natural language lost critical context at each handoff, compounding errors across multiple agent interactions.
Agents think in milliseconds, legacy infrastructure doesn't. LinkedIn, Walmart and Zendesk shared how they closed the gap at VB Transform 2026
Enterprise infrastructure designed for human workflows became the bottleneck when scaling AI agents to production, not the underlying AI models themselves. Companies like LinkedIn, Walmart, and Zendesk each encountered different infrastructure failures when moving agents from pilot to production environments.
Brex built its AI agent policy by watching what agents actually do, not by writing rules first
Brex built CrabTrap, an open-source HTTP/HTTPS proxy that intercepts agent network traffic and uses a large language model to approve or deny requests based on policy rules. The platform monitors every API call agents make at the network layer rather than relying on traditional SDK-level permissions or model guardrails.
How Apple’s big lawsuit could disrupt OpenAI’s IPO plans
Apple filed a trade secrets lawsuit against OpenAI, alleging misconduct by its chief hardware officer and claiming over 400 former Apple employees work there. The lawsuit's timing potentially complicates OpenAI's reported plans to pursue an initial public offering.
Bunkerhill raises $55M to scale agentic AI across health systems
Bunkerhill Health raised $55 million in Series B funding to expand its Carebricks agentic AI platform across health systems. The round included participation from Sequoia Capital, Felicis, Optum Ventures, and Y Combinator.
Patreon stops asking AI bots not to scrape — and starts blocking them
Patreon partnered with Cloudflare to actively block AI bots from scraping creator content, shifting away from passive robots.txt files. The platform is now preventing unauthorized data collection for AI model training through active technical defenses.
NVIDIA AI Releases Nemotron 3 Embed: An Open Embedding Collection Whose 8B Checkpoint Ranks #1 on RTEB
NVIDIA released three open embedding model checkpoints including an 8B version ranking first on RTEB with a 78.46 average NDCG@10 score. The 1B model was created through pruning and distillation from the 8B teacher, while the NVFP4 variant maintains 99% plus retrieval accuracy at up to 2x Blackwell throughput.
China’s Moonshot AI releases Kimi K3, the largest open-source model ever, rivaling top U.S. systems
Moonshot AI released Kimi K3, a 2.8-trillion-parameter open-source model described as the world's largest, with performance comparable to top U.S. proprietary systems from Anthropic and OpenAI. Full model weights are scheduled for release on July 27, and the model is currently accessible for testing at kimi.com.
The AI compute gap: Enterprises are buying infrastructure faster than they can measure what it costs
Across 107 enterprises, AI infrastructure spending is accelerating faster than organizations can track its costs, with GPUs operating at half utilization or less. Fewer than half of enterprises rigorously measure what their compute actually costs, creating a gap between rapid investment and economic visibility.
The agent security gap: 54% of enterprises have already had an AI agent incident, and most still let agents share credentials
Fifty-four percent of enterprises have experienced confirmed AI agent security incidents, while most still allow agents to share credentials rather than isolating them with individual identities. Only 30 percent of organizations isolate their highest-risk agents, and security controls remain borrowed from model providers instead of purpose-built for agent management.
Google Vids now lets you star in your own AI videos
Google Vids now allows users to create AI avatars that represent themselves in videos. The feature uses Gemini Omni-powered tools to generate and edit videos from prompts and reference images.
The AI context gap: Enterprise AI organizations have a trust problem, not a retrieval problem — and most are still building the fix
A majority of 57% of enterprises report their AI agents have produced confident but incorrect answers traced to missing or inconsistent business context in the past six months. Most enterprises are building a governed semantic layer to address this context gap, while provider-native retrieval has overtaken dedicated vector databases as the default source feeding AI agents their business information.