Archive

377 stories in AI

Monday, 20 July 2026

At VB Transform 2026, Zillow's engineering chief said AI ROI numbers only hold up if you measure before you build

Zillow built an AI architecture designed to maintain customer context across multiple touchpoints in real estate transactions rather than relying on isolated chatbots. The company discovered that establishing a persistent context layer proved more challenging than organizing its data infrastructure, despite initially focusing resources on data foundation and governance.

The cleanup trap: Stop asking RAG to fix bad data

Organizations mistakenly believe retrieval-augmented generation systems can fix fragmented legacy data at the retrieval layer, but production AI failures typically stem from inadequate enterprise data foundations rather than model limitations. The cleanup trap occurs when companies pipe ungoverned data into LLM orchestrators expecting the retrieval layer to patch problems through vector databases and embedding pipelines.

Safety guardrails blocked Hugging Face's defenders, not the attacker, when an AI agent breached its systems

Safety guardrails on frontier AI models blocked Hugging Face's incident response team from analyzing exploit data during a breach, while an autonomous AI agent moved undetected across the company's infrastructure for a weekend. Commercial models cannot cryptographically or organizationally distinguish between legitimate incident responders and malware authors requesting analysis.

The Download: AI hiring biases, and weather data sabotage

AI systems are more likely than humans to develop biases during hiring decisions. Résumés may be screened by artificial intelligence before human reviewers evaluate them.

Bristol Myers Squibb Building Life Science Industry’s Most Advanced AI Factory on NVIDIA Vera Rubin

Bristol Myers Squibb is deploying a second NVIDIA DGX SuperPOD to expand its AI computing capacity for drug discovery and development work. The company already operates one of the largest AI clusters in life sciences and is doubling its computational infrastructure investment.

US public health agencies to test OpenAI and Anthropic AI models

Ten state, local, tribal, or territorial jurisdictions will test generative AI models from OpenAI and Anthropic through a new program called PULSE. The initiative involves the Coalition for Health AI and Accenture to support public health departments in evaluating these AI tools.

Kimi K3 open-weight model: China’s biggest AI is a bet on memory, not compute

Moonshot AI released Kimi K3, an open-weight model with 2.8 trillion parameters, making it the largest open-weight model released to date. The model launched on July 16 and represents what the industry classifies as the 3T tier bracket.

AI is more likely than humans to form biases when hiring

Large language models develop biases when screening job applications, picking up prejudices from training data and creating additional ones independently. A study found AI systems show greater hiring discrimination than humans across multiple demographic categories in résumé evaluation.

AI confidence just dropped 17 points in six months. That’s actually great news.

IT leaders' assessment of their organizations' AI maturity fell from 40% to 23% over six months. Organizations deploying AI agents into production are revising downward because they encounter real-world operational challenges that pilots do not reveal.

Best Local LLMs You Can Run on a Single 24GB GPU in 2026: Qwen, Gemma, Mistral, DeepSeek Compared

Six open-weight models fit on a single 24GB GPU at Q4_K_M quantization, including Qwen3.6, Gemma 4, Mistral Small, gpt-oss-20b, and DeepSeek-R1-Distill. The guide evaluates each model's VRAM requirements, licensing terms, and performance strengths for local inference.

Sunday, 19 July 2026

Feyn AI Releases SQRL, a Text-to-SQL Model Family That Inspects the Database Before Writing a Query

SQRL, a text-to-SQL model family, inspects databases with read-only probes before generating queries. The flagship SQRL-35B-A3B achieves 70.6% execution accuracy on BIRD Dev and distills into smaller 4B and 9B self-hostable versions.

Alibaba Previews Qwen3.8-Max, a 2.4 Trillion-Parameter Multimodal Model, Days After Moonshot’s Kimi K3 Open-Weight Launch

Alibaba previewed Qwen3.8-Max, a 2.4 trillion-parameter multimodal model available at 10% standard pricing on multiple platforms. The preview lacks published benchmark tables, model cards, licenses, per-token pricing, or active-parameter counts.

Perplexity AI Releases WANDR: An Open Benchmark Evaluating Research Agents That Must Search Wide And Deep

An open benchmark named WANDR containing 500 evidence-heavy tasks evaluates whether research agents can discover multiple qualifying entities with cited, re-verifiable evidence. Perplexity Search as Code achieved the highest scores at 0.363 soft F1 and 0.133 hard F1 on the evaluation harness.

Kimi K3 vs DeepSeek V4 Pro vs GLM-5.2: Open Trillion-Scale MoE Models Compared on Benchmarks, License, and Serving Cost

Three trillion-scale mixture-of-experts models with open weights were compared across benchmarks, licensing terms, and serving costs. The comparison evaluated Kimi K3, DeepSeek V4 Pro, and GLM-5.2 using measured intelligence scores and practical deployment expenses.

Friday, 17 July 2026

Capital One releases VulnHunter, an open-source AI tool that finds software flaws before hackers do

An open-source AI tool called VulnHunter scans source code for exploitable vulnerabilities and proposes fixes before deployment. The tool uses "attacker-first forward analysis" to trace how adversaries would enter systems through APIs, network messages, or file uploads.

Intuit scrapped its own AI agent architecture twice in four months. At VB Transform 2026, its AI VP called that the fast path

Intuit rebuilt its AI agent architecture twice within four months, first consolidating specialist agents under a central orchestrator, then switching to a skills-based system after the orchestrator collapsed from complexity. The orchestrator failed because agents passing results in natural language lost critical context at each handoff, compounding errors across multiple agent interactions.

Agents think in milliseconds, legacy infrastructure doesn't. LinkedIn, Walmart and Zendesk shared how they closed the gap at VB Transform 2026

Enterprise infrastructure designed for human workflows became the bottleneck when scaling AI agents to production, not the underlying AI models themselves. Companies like LinkedIn, Walmart, and Zendesk each encountered different infrastructure failures when moving agents from pilot to production environments.

Brex built its AI agent policy by watching what agents actually do, not by writing rules first

Brex built CrabTrap, an open-source HTTP/HTTPS proxy that intercepts agent network traffic and uses a large language model to approve or deny requests based on policy rules. The platform monitors every API call agents make at the network layer rather than relying on traditional SDK-level permissions or model guardrails.

How Apple’s big lawsuit could disrupt OpenAI’s IPO plans

Apple filed a trade secrets lawsuit against OpenAI, alleging misconduct by its chief hardware officer and claiming over 400 former Apple employees work there. The lawsuit's timing potentially complicates OpenAI's reported plans to pursue an initial public offering.

Bunkerhill raises $55M to scale agentic AI across health systems

Bunkerhill Health raised $55 million in Series B funding to expand its Carebricks agentic AI platform across health systems. The round included participation from Sequoia Capital, Felicis, Optum Ventures, and Y Combinator.

Patreon stops asking AI bots not to scrape — and starts blocking them

Patreon partnered with Cloudflare to actively block AI bots from scraping creator content, shifting away from passive robots.txt files. The platform is now preventing unauthorized data collection for AI model training through active technical defenses.

NVIDIA AI Releases Nemotron 3 Embed: An Open Embedding Collection Whose 8B Checkpoint Ranks #1 on RTEB

NVIDIA released three open embedding model checkpoints including an 8B version ranking first on RTEB with a 78.46 average NDCG@10 score. The 1B model was created through pruning and distillation from the 8B teacher, while the NVFP4 variant maintains 99% plus retrieval accuracy at up to 2x Blackwell throughput.

Thursday, 16 July 2026

China’s Moonshot AI releases Kimi K3, the largest open-source model ever, rivaling top U.S. systems

Moonshot AI released Kimi K3, a 2.8-trillion-parameter open-source model described as the world's largest, with performance comparable to top U.S. proprietary systems from Anthropic and OpenAI. Full model weights are scheduled for release on July 27, and the model is currently accessible for testing at kimi.com.

The AI compute gap: Enterprises are buying infrastructure faster than they can measure what it costs

Across 107 enterprises, AI infrastructure spending is accelerating faster than organizations can track its costs, with GPUs operating at half utilization or less. Fewer than half of enterprises rigorously measure what their compute actually costs, creating a gap between rapid investment and economic visibility.

The agent security gap: 54% of enterprises have already had an AI agent incident, and most still let agents share credentials

Fifty-four percent of enterprises have experienced confirmed AI agent security incidents, while most still allow agents to share credentials rather than isolating them with individual identities. Only 30 percent of organizations isolate their highest-risk agents, and security controls remain borrowed from model providers instead of purpose-built for agent management.

Google Vids now lets you star in your own AI videos

Google Vids now allows users to create AI avatars that represent themselves in videos. The feature uses Gemini Omni-powered tools to generate and edit videos from prompts and reference images.

The AI context gap: Enterprise AI organizations have a trust problem, not a retrieval problem — and most are still building the fix

A majority of 57% of enterprises report their AI agents have produced confident but incorrect answers traced to missing or inconsistent business context in the past six months. Most enterprises are building a governed semantic layer to address this context gap, while provider-native retrieval has overtaken dedicated vector databases as the default source feeding AI agents their business information.

← Prev178913Next →

Get feedd. daily

Top stories in your inbox every morning. Pick what you want.

No spam. Unsubscribe anytime.