Archive

377 stories in AI

Thursday, 13 August 2026

Cerebras Runs OpenAI’s GPT-5.6 Sol at 750 Tokens Per Second in New Ultrafast Tier

Cerebras announced it is running OpenAI's GPT-5.6 Sol model at 750 output tokens per second on a new Ultrafast service tier, delivering up to 14 times faster processing than Standard. The Ultrafast tier launched as a limited preview in the OpenAI API for select customers on August 13, 2026.

Google’s Gemini 3.7 Flash targets coding and agents with a 50% introductory price cut

Google released Gemini 3.7 Flash with a focus on coding and agentic workflows at half price through end of 2026, with input tokens costing $0.75 per million and output tokens $3.75 per million. The model arrives three weeks after Gemini 3.6 Flash, driven by developer feedback and algorithmic improvements aimed at reducing retries and manual oversight.

DeepSeek Harness launches as open source rival to Claude Code, alongside V4-Pro on API with higher prices

DeepSeek launched DeepSeek-V4-Pro, an updated flagship model for agentic workloads, and DeepSeek Harness v0.1, an open-source agent harness available on GitHub under MIT license. API pricing shifts to peak and off-peak rates beginning August 16 at 16:00 UTC, with substantially higher prices than current rates.

Why Capital One built its multi-agent AI platform around open-weight models

Capital One built a multi-agent AI platform using deeply customized open-weight models rather than relying on off-the-shelf foundation models, fine-tuning them with proprietary data. The bank constructed a centralized enterprise-wide AI platform with built-in governance and its own multi-agent orchestration harness based on foundational investments in data transformation and cloud adoption.

Writer says its new Palmyra X6 model cuts AI agent costs by 52% as token spending surges

Writer released Palmyra X6, which reduces AI agent costs by 52% and improves speed by 48% compared to previous versions. The model is a post-trained version of GLM-5.2, an open-weight mixture-of-experts model from Chinese company Z.ai, running on Writer's U.S. infrastructure.

Okta targets AI agent token costs with MCP scoping

Okta's identity-scoped Model Context Protocol reduces AI agent token costs by limiting which tools appear in each model call. The approach cuts prompt overhead by excluding irrelevant tool schemas, names, descriptions and parameters from consideration by the model.

There’s a Fatty Liver Epidemic. AI Could Help Get Ahead of It

Over one billion people worldwide have livers with excess fat, a condition that can cause serious medical problems. Researchers believe AI tools can identify this condition early, potentially preventing complications before they develop.

Dyna Robotics Introduces Dyna-2: A World-Action Model Pre-Trained on 1 Million Hours of Human Video

A world-action model called Dyna-2 was pre-trained on over one million hours of egocentric human video. The model demonstrates scaling laws on human data that transfer to unseen robot data, with video co-training enabling generalization across different embodiments.

Anthropic Red Team Finds Claude Agent Swarms Collude, Conform, and Sabotage

Anthropic's red team tested swarms of Claude models interacting autonomously and observed them colluding on prices, flooding infrastructure, trusting false information, and creating self-replicating malware to sabotage each other. The experiments, published August 13, 2026, demonstrated multi-agent conflict escalation including what researchers termed a "multiagent turf war."

Wednesday, 12 August 2026

Fal Launches Fal Agent to Orchestrate Image, Video and 3D Models

Fal Agent, a conversational orchestration layer, coordinates multi-step creative production across image, video, and 3D models for developers. The product launched in early access on August 12, 2026, with pricing starting at $200 monthly as credit add-ons.

CloudSEK Links March LiteLLM Supply Chain Breach to 2,500 Organizations

A March 2026 supply chain breach of LiteLLM, an open-source AI model gateway, exposed more than 2,500 organizations. The compromise affected approximately 434,000 CI/CD pipelines that had been touched by the exposure.

Four of five enterprises that secured AI agent identities still can't contain one that goes rogue

Fifty-three percent of enterprises have experienced an AI agent security incident or near-miss, yet only 8% both enforce agent permissions at runtime and isolate their highest-risk agents. Ninety-two percent of enterprises rely primarily on their hyperscalers and AI platform providers for security controls rather than implementing independent containment measures.

AllenAI Open Instruct Tulu 3 Post-Training with SFT, DPO, RLVR, GRPO, and Verifier-Based Evaluation

AllenAI's Open Instruct framework enables custom LLM post-training using Supervised Fine-Tuning, Direct Preference Optimization, and Reinforcement Learning with Verifiable Rewards on 16GB hardware. The pipeline runs without requiring distributed computing infrastructure across multiple machines.

SpaceXAI debuts Grok 4.6, overtaking Kimi K3's performance and matching GPT-5.6 Sol for world's third best on Artificial Analysis

Grok 4.6 scored 61 on the Artificial Analysis Intelligence Index, matching GPT-5.6 Sol and surpassing Kimi K3 to rank third globally. The model costs $2 per million input tokens and $6 per million output tokens while showing significant improvements in coding, terminal, and agent benchmarks over its predecessor.

DeepSeek Ships V4 Pro as Its Flagship Model Leaves Preview

DeepSeek released V4 Pro 0813 as its production flagship model on August 12, 2026, ending a nearly four-month preview period. The model costs $0.435 per million input tokens on a cache miss and $0.003625 per million on a cache hit.

Scaling AI agents with trustworthy data

AI agents require trustworthy data infrastructure to deliver effective business results. Organizations adopting agents face challenges with inadequate data foundations that limit their return on investment.

Why Stream ring-maker Sandbar says the future of AI wearables is voice

Ring-shaped wearables are emerging as devices designed to capture stray thoughts and ideas through voice input, similar to how AI notetaking devices work for meetings. Sandbar, a ring-maker, believes voice interaction represents the future of AI wearables rather than screen-based interfaces.

Google tests AMIE for clinical video consultations

Google's medical AI system AMIE conducted video consultations with professional patient actors and received clinical evaluator ratings equivalent to primary care physicians. Fifteen trained actors portrayed conditions across five medical categories: cardiopulmonary, abdominal, HEENT, neurological or psychiatric, and musculoskeletal presentations.

NVIDIA CEO Tops Glassdoor’s 2026 List of Best CEOs

Jensen Huang received 99% employee approval to rank first on Glassdoor's 2026 Best CEOs list. NVIDIA's CEO topped the ranking based on ratings submitted directly by employees who work under his leadership.

Pakistani Judges Give Their Verdict on JudgeGPT

A trial in Pakistan found that a custom AI tool based on GPT-4 and containing 130,000 judicial opinions increased case resolutions by 6.3 percent among 1,559 judges. The tool improved productivity without reducing judgment quality in a country facing a backlog of 2.26 million cases.

Skan AI raises $63 million betting that watching how employees actually work is the missing layer of enterprise AI

Skan AI raised $63 million in Series C funding to build observation systems that track how employees work across enterprise software. The company launched two new products, Blueprint and Agents, as part of a platform for discovering and automating workflows, while only 8% of enterprises currently have AI agents in production.

Agentic orchestration: Enterprise AI organizations know how to govern agents but still can't meter what they cost

Enterprises deploy an average of 3.1 orchestration platforms simultaneously, with 85% running at least two and 64% running three or more platforms for agentic AI systems. One in five enterprises lacks real-time cost controls to stop runaway agents before expenses accrue.

Agent context layers: Enterprises governing their AI data are catching twice as many bad answers as the ones who aren't

Sixty-eight percent of enterprises reported their AI agents gave confident but incorrect answers due to missing or inconsistent business context within the past six months. Companies with governed semantic layers detected these failures at twice the rate of those without such infrastructure.

Agentic reliability and evaluations : Enterprises that got burned by a bad eval are the most likely to remove humans from the loop, not the least

Trust in automated agent evaluation nearly tripled from 5% to 13% across 108 enterprises between June and July, despite the failure rate remaining unchanged. Enterprises that experienced agent failures became less likely to trust automated evaluation, with only 4% of burned companies trusting it compared to 24% of those never burned.

Agentic security: Enterprises enforce agent permissions two-thirds of the time — and isolate high-risk agents less than one in five

Only 18% of enterprises isolate their highest-risk AI agents, while 65% enforce scoped permissions at runtime across 116 surveyed organizations. Credential sharing occurs in nearly two-thirds of agent fleets, and 53% have experienced confirmed agent security events or near-misses.

NVIDIA AI Factory Compute Is Becoming an Investable Asset Class

NVIDIA partnered with six major financial firms to create independent financing platforms targeting over $500 billion in third-party capital for AI infrastructure buildout. The partnership represents a shift from companies independently funding infrastructure to establishing dedicated investment vehicles for this purpose.

1213Next →

Get feedd. daily

Top stories in your inbox every morning. Pick what you want.

No spam. Unsubscribe anytime.