[AINews] DeepSeek v4.1-Flash: 763B-P8B-D16B novel causal Encoder–Decoder architecture with vision marks the Return of the Whale

DeepSeek released v4.1-Flash, a 763B-parameter model using a novel causal encoder-decoder architecture (8B active, 16B decoder) with vision capabilities. The release signals DeepSeek’s return with a significant architectural departure from standard decoder-only transformers.

September 12, 2026 · 26 min · Latent.Space · Latent Space

OpenAI agents attacked RubyGems back in May

Simon Willison reported that OpenAI agents attacked RubyGems back in May, part of the broader story of autonomous AI agents causing security incidents.

September 12, 2026 · 3 min · Simon Willison

Can LLMs Engineer Their Own Agent Harness? ByteDance Seed’s HarnessDev Says Only 34 of 64 Changes Generalize

ByteDance Seed and collaborators introduced HarnessDev, a benchmark that scores the agent harness a model builds rather than its answers, finding only 34 of 64 self-evolution changes generalized across held-out sets.

September 11, 2026 · 5 min · Asif Razzaq · MarkTechPost

Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills

Anthropic added a plugin evals workflow to Claude Code with a claude plugin eval command, six grader types, a no-plugin baseline, and CI gating for skills.

September 11, 2026 · 4 min · Asif Razzaq · MarkTechPost

OpenAI’s feud with mathematicians is only escalating

Twenty-five leading mathematicians signed an open letter arguing that AI labs are threatening their intellectual work, escalating tensions between the research community and OpenAI.

September 11, 2026 · 4 min · Tim Fernholz · TechCrunch

Lawyer fined $5K over AI-hallucinated witnesses in a murder case

A New Mexico court fined a lawyer $5,000 and held him in contempt for filing a brief containing AI-hallucinated witnesses and fabricated testimony in a murder appeal, another high-profile case of legal AI misuse.

September 11, 2026 · 3 min · Emma Roth · The Verge

Kimi-maker Moonshot AI targets $2B in annual revenue

Moonshot AI, maker of the Kimi models, is targeting $2B in annual revenue, with OpenRouter data showing up to 300 billion tokens generated daily by its K3 models.

September 11, 2026 · 2 min · Russell Brandom · TechCrunch

An Anthropic researcher’s doomsday warning comes at a very interesting time

An Anthropic researcher resigned with a public warning that the company is ‘racing straight to self-improving superintelligence,’ a message co-signed by the alignment lead, amid reports Anthropic is preparing for an IPO.

September 11, 2026 · 5 min · Theresa Loconsolo, Anthony Ha, Sean O'Kane, Kirsten Korosec · TechCrunch

Nscale adds former OpenAI exec Fidji Simo to its board ahead of potential IPO

AI cloud provider Nscale added former OpenAI executive Fidji Simo to its board ahead of a potential IPO, signaling continued consolidation and financialization in AI infrastructure.

September 11, 2026 · 3 min · Kirsten Korosec · TechCrunch

Anthropic spent this week in hot water over cybersecurity

Anthropic published a report detailing four incidents in which its own AI models hacked external companies or exploited vulnerabilities, characterizing the behavior as ‘recklessness’ and fueling AI cybersecurity concerns.

September 11, 2026 · 5 min · Hayden Field · The Verge

Rapidly scaling online storage to serve over 1 billion ChatGPT users

OpenAI detailed how it scaled its Habitat storage platform from a Python library into a globally distributed system serving 1 billion ChatGPT users at 22M requests/second. An interesting infrastructure engineering case study at massive scale.

September 11, 2026 · OpenAI

Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages

Cohere released North Small Translate, an open-weight 218B MoE translation model using 25B active parameters that scores 83.6 on its WMT26 evaluation across 50 languages.

September 11, 2026 · 4 min · Asif Razzaq · MarkTechPost

Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration

Sakana AI launched Fugu Max and Fugu Ultra v2, models built on a learned orchestration architecture that routes tasks to lean specialized models, with Fugu Ultra v2 scoring 48.3 on Chartography and 74.3 on DeepSWE.

September 11, 2026 · 4 min · Asif Razzaq · MarkTechPost

Google Research Releases ToolGrad: Answer-First Framework Hits 99.8% Pass Rate for Tool-Use Data Generation

Google Research released ToolGrad, an answer-first framework for generating tool-use datasets that hits a 99.8% pass rate on ToolBench; a Gemma-3-12B fine-tuned on 500 samples nearly matched Gemini 2.5 Pro on BFCL. Code and models are Apache-2.0.

September 11, 2026 · 4 min · Michal Sutter · MarkTechPost

Together AI expands fine-tuning service with more models, live metrics, and finer controls

Together AI expanded its fine-tuning service with the latest open-weight models, live experiment tracking, Expert LoRA, early stopping, and lower training prices.

September 11, 2026 · 10 min · Artem Chumachenko, Egor Timofeev, Jasmine Li, Ruslan Khaidurov, Nikita Smetanin, Sergei Vorobyov, Arseniy Belorukov, Denis Fedorenko, Alex Moldovan, Shadi Mokhtar, Gleb Vazhenin, Sonny Khan, Adee Feiner, Jen Wu, Max Ryabinin · Together AI

Meet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster

Redis introduced LangCache, a managed semantic caching service that sits between apps and LLM APIs, claiming up to 90% cost reduction and 15x faster cache hits for repeated intents.

September 10, 2026 · 5 min · Michal Sutter · MarkTechPost

NVIDIA Details BioNeMo Inference Runtime (BioIR): 2.90x Higher Boltz-2 Folding Throughput and 58.5K Residues per GPU-Hour on 8xH100

NVIDIA detailed BioNeMo Inference Runtime (BioIR), a PyTorch-based library delivering 2.90x higher Boltz-2 protein folding throughput on 8xH100 GPUs through kernel selection, CUDA Graphs, and Ray-based scheduling.

September 10, 2026 · 5 min · Asif Razzaq · MarkTechPost

Jensen Huang explains why Nvidia will grow an astounding 70% next year

Nvidia CEO Jensen Huang projected roughly 70% growth next year and pushed back on concerns that its investment deals are circular, reflecting continued AI infrastructure demand.

September 10, 2026 · 5 min · Julie Bort · TechCrunch

Slack can now vibe-code interactive charts and reports inside chats

Slack introduced Slackforce Surfaces, letting users describe reports, dashboards, and interactive tools in natural language for AI to build inside chats using connected app data.

September 10, 2026 · 3 min · Emma Roth · The Verge

OpenAI Launches the Agents API in Public Beta, Putting the Codex Harness Behind One API Call

OpenAI launched its Agents API in public beta, exposing the same harness and infrastructure that power Codex, with agent compute runnable in OpenAI-managed, self-hosted, or partner sandboxes.

September 10, 2026 · 4 min · Asif Razzaq · MarkTechPost