EAIDaily — September 13, 2026

English AI Daily Report focusing on AI Coding and Embodied Intelligence

EAIDaily — 2026-09-13

AI coding + embodied intelligence daily brief. Today: harness commoditization meets pacing-the-frontier as the Anthropic/Nvidia IPO narrative lands; DeepSeek V4.1-Flash redefines the inference cost floor; Sakana’s Fugu Ultra v2 and Devin Fusion push multi-model orchestration; XPeng IRON crosses mass-production; China’s embodied-AI data-loop bubble pops into a public fight.


1. Anthropic & Nvidia: $100B IPO talk at ~$2T valuation, while Dario calls to “pace the frontier”

  • What happened: Reuters reports Nvidia is in talks to invest up to $10 billion as an anchor in Anthropic’s planned mega-IPO, which could raise up to $100 billion at a ~$2 trillion valuation. On September 12, CEO Dario Amodei published the essay “We Must Pace the Frontier” urging a coordinated slowdown in frontier model releases, paired with permanent employee-level access for third-party evaluators. A widely circulated screenshot shows Sam Altman publicly agreeing, while in a Fortune interview Altman separately said an OpenAI IPO in 2026 would be “ill-advised.”
  • Why it matters: For the first time the two largest model labs are openly converging on a narrative of restraint even as capital velocity accelerates (Nvidia-as-customer-and-investor stack). “Pacing” plus third-party evaluator access is being floated as a governance substrate that, if adopted, would lock the current frontier leaders in and raise the bar for new entrants — a self-serving safety story critics (Sequoia’s tszzl) frame as regulatory capture.

2. OpenAI Agents API in public beta — harness becomes a public utility

  • What happened: OpenAI opened its Agents API to all developers on September 10. The product exposes the same harness Codex/ChatGPT-for-Work runs on internally (Agent / Environment / Session / Events), with native context compaction, tool search, programmatic tool calls, multi_agent sub-agents, and a choice of nine sandbox partners (Blaxel, Cloudflare, Daytona, DigitalOcean, E2B, Modal, Oracle, Runloop, Vercel). No extra fee — billing is token + compute. Early customer numbers: Ciridae 0.71→0.85 score, ~4× lower latency; SafetyKit ~60% lower per-case cost; Hypha 86% fewer agent failures.
  • Why it matters: This is the moment the agent harness leaves the model and becomes a managed service. On the same day, X’s coders ranked “Harness Engineering vs Claude Code / Codex workflow” the top topic (4,470 likes, 166k+ views), and multiple long-reads (Anthropic workshops, Pi Agent) argued that the harness layer — not the LLM — is now the binding constraint on Terminal-Bench scores. Combined with Devin Fusion (story 4), this is the clearest signal yet that the moat has migrated from model to orchestration.

3. DeepSeek V4.1-Flash: 552B MoE at 890 bytes/token — sets new cost floor for long-running loops

  • What happened: DeepSeek launched V4.1-Flash — a 552B-parameter (763B-P8B-D16B) MoE with native visual understanding, using CED (halves prefill compute), CSA2 (3D KV compression), FP4 quantization, and a Bounded Replay training scheme. Headline number: global KV cache compressed to 890 bytes/token with 1M context support at ~$0.27 per task, scoring 40 on Artificial Analysis benchmarks — surpassing the larger V4 Pro. Practical price test: a cinema landing page cost $1.21 on Claude Fable 5 vs $0.026 on V4.1-Flash (≈1/40th the price, similar quality).
  • Why it matters: Combined with DeepSeek’s grey-test launch today of an AI voice conversation feature (four voices), the model compresses the per-token economics of an agent loop to the point where running your own weights on rented H200s (~$13 vs ~$97 on Claude Sonnet 5 for 1k-doc batch jobs) becomes the obvious choice for cost-sensitive teams. The “agents are compute-cheap” era just arrived in production — and Loop Daily flags the cost floor as the binding metric going forward.

4. Cognition’s Devin Fusion + Sakana’s Fugu Ultra v2 — multi-model orchestration hardens

  • What happened: Cognition launched Devin Fusion, pairing a frontier lead model (Claude Fable 5.1 or GPT-6 Astra) with a cheap sidekick (in-house SWE-2, GLM, including free options), scoring 62 on Artificial Analysis — the first multi-model coding agent benchmarked on the index, hitting that score at ~36% lower cost than Claude Code 62.2. Separately, Sakana AI shipped Fugu Ultra v2, the second iteration of its multi-model orchestration flagship with claims of world-class coding performance. Meta’s Muse Spark drew strong early praise (Sourcegraph CEO called it “actually pretty pleasant to work with”; a prominent tools reviewer named it the “best agent implementation so far” — easy, intuitive, free).
  • Why it matters: This is the second day in a row Cognition pushes the “configurable effort” + dual-model planner/executor pattern into the leaderboard conversation; combined with Sakana’s orchestration-first posture and Muse’s free-tier pricing, the competitive narrative has shifted from “which model wins” to “which model + which sidekick + which harness wins per benchmark.”

5. XPeng IRON starts mass production; UBTECH inks Egypt exclusive; Hefei deploys 19 humanoid patrol robots

  • What happened: At IFA Berlin 2026, XPeng called the start of IRON mass production following the September 8 first unit rolling itself off the line. Production spec: ~1.7 m tall, ~65 kg, 76 DoF (21 in each hand), built in an existing factory (volume undisclosed). Target use case: household chores (floor sweeping, folding laundry). On September 11, UBTECH signed Misr Elsalam for Development & Advanced Technology (Egypt) as its exclusive regional agent for education / business / public affairs. In Hefei, 19 humanoid robots patrol on a dedicated 5G-A network — China’s first communications plan specifically designed for embodied robots. Separately, ShengShu Technology’s Motus2 world model hit 84% average success on five real-robot dexterous manipulation tasks; Beijing’s Shijingshan district unveiled a nationwide-first “unmanned post office + cultural-creative retail” humanoid robot terminal (±0.03 mm placement, full order-to-delivery autonomous loop).
  • Why it matters: Three distinct routes to scale in one day: (a) automotive-grade humanoid line with self-rolling-off (XPeng), (b) overseas channel capture at 2–4× export price arbitrage (UBTECH), and (c) infrastructure-layer (5G-A + world models) becoming the government’s preferred moat. Note also that Motus2 explicitly closes the loop — embedding an evaluator inside the world model — which is a parallel to today’s “harness engineering” theme in coding.

6. Galbot files police report as China’s embodied-AI “data-collection loop” goes public

  • What happened: On September 12, Chinese embodied-AI startup Galbot issued a strongly worded statement saying online content attacking the company and its leadership is false, that it has filed a police report, and will pursue legal action. The trigger: a September 10 WeChat Moments post by Mech-Mind founder/CEO Shao Tianlan that named Galbot directly in a critique of “consortium-style entrepreneurship” — using data-collection centres, leasing companies, and related-party transactions to manufacture unsustainable revenue. Mechanically (per Lingyu Zhineng CEO Jin Ge): a ~¥150k robot, ¥120–150/hour data buy-back, 100-unit centre producing 400 hours/day, robot pays for itself in 1 year; in another variant, embodied-AI companies set up JV data-collection firms with local governments, the JV funds robot purchases, and the embodied-AI firm buys the data back. One investor summarised this as a 5-year local-government “bond” repaid in data-purchase instalments, inflating both revenue and valuation. Broader numbers: 288 funding deals in Chinese robotics H1 2026 (¥46B+ disclosed, surpassing all of 2025); 15+ large-scale data-collection sites nationally. Secondary market moved opposite: Unitree closed Sep 11 at ¥477/share (¥193B market cap, down from ¥444.9B peak on Aug 19 listing day); Mech-Mind broke issue price on Sep 1 and closed Sep 11 at HK$95 (issue HK$101.7).
  • Why it matters: The first publicly contested case of real-robot teleoperation data scarcity translating into a revenue-inflation pattern. The takeaway is structural: until the data problem is solved by cheaper sim-to-real or world-model-based proxies (see Motus2 story above), the commercial incentives will keep rewarding circular revenue over genuine product-market fit — and the secondary market is already discounting.

7. Lexoo Technology alleges OpenAI copied its Aether model; launches litigation

  • What happened: Suzhou-based Lexoo Technology CEO Guo Renjie published an open letter to OpenAI point-by-point comparing his company’s Aether large model to OpenAI’s recent announcements. Claimed parallels: recursive self-improvement, AI for compute optimisation, splitting alignment into goal + value alignment, brain–cerebellum–cortex–spinal-cord-inspired architecture. Quote (CNBC translation): “People always say big tech companies have intelligence networks monitoring the entire internet — this time I believe it. This is a direct, unmodified distillation of us.” Litigation announced.
  • Why it matters: First publicly-litigated Chinese-vs-OpenAI distillation accusation at the architectural-idea level (not just data). The four specific overlaps named — recursive self-improvement, AI-for-compute, split alignment, neuroscience-inspired architecture — are exactly the four motifs that have spread across labs in 2026 (Anthropic, DeepSeek, Sakana, the Zhejiang SPIRE brain+cerebellum design). Distinguishing “convergent design” from “unauthorised distillation” is the next IP battleground in frontier AI.

8. Sundries shaping the context

  • OpenAI launches ChatGPT for Financial Services with Morgan Stanley + Evercore, wired into Daloopa / PitchBook / LSEG News, with source-tracing for every figure (verified-citation matters more than “knowing a lot” for sell-side compliance).
  • OpenAI ships Habitat internal storage write-up — claims >1 billion weekly users across ChatGPT + API, tens of millions of RPS.
  • César de la Fuente lab (Penn) using Codex/ChatGPT to mine genomes of living + extinct organisms for novel antimicrobial candidates — concrete codex-as-science win against treatment-resistant microbes.
  • Z.AI launches $5B financing operation ($2B HK placement + ~$3B convertible), expanding Chinese compute capacity alongside a Cohere up-to-$3B round at ~$20B (post Aleph Alpha merger, with Canadian + German government participation) — Canadian/German “AI sovereignty” via direct capitalisation, not just infrastructure subsidies.
  • California signs 13 laws restricting social media + AI chatbots for under-16s (no infinite scroll / autoplay, mandated AI safety assessments) — first U.S. state to bundle social media + chatbot AI as the same regulatory object.
  • 25 Fields Medalists publish “A Severe Misalignment of AI in Mathematics” (Tao, Scholze, Deng, etc.) — first open academic pushback from the math community; Tao separately warns that if pure math becomes “hobby chess,” it dies; another widely-shared take argues the Millennium Prize problems are now Goodharted metrics labs are over-fitting.
  • DeepSeek grey-tests AI voice conversation with four voices in the app (Chinese consumer AI voice race heats up — Tencent Yuanbao, Kimi, Doubao all live).
  • Anthropic Threat Intelligence Report (Sep 2026): Russian state media used Claude to “transform” Romanian/Moldovan news via a former Sputnik Moldova editor’s workflow — including fabricated content targeting President Maia Sandu ahead of the 2025 parliamentary elections — first documented case of Claude inside a state propaganda production chain.
  • Token demand: China Telecom Research Institute projects 2026 China token consumption ≈ 100 trillion (10亿亿), growing to >350 quadrillion (3500亿亿) by 2030 (CAGR ~12×) — agent-scale deployment is the binding signal.
  • Anthropic Threat Intelligence Report also documents: Claude used to build software for kamikaze drones in the Ukraine war — first public attribution of a frontier commercial model inside an autonomous weapons pipeline.
  • Yoshua Bengio publishes “Why are AI agents lying, cheating and coordinating?” — the new agent-safety sub-literature now has an organiser-in-chief.
  • Nvidia reportedly in talks to invest in Anthropic mega-IPO (see story 1) — supplier-as-anchor-investor pattern consolidates.

Trend of the day: The orchestration layer has won. Harness engineering, dual-model planners, world models with embedded evaluators, 5G-A for embodied agents, NVMe-shaped IPOs — across both AI coding (OpenAI Agents API, Devin Fusion, Muse Spark) and embodied intelligence (XPeng IRON, Motus2 with embedded evaluator, Hefei’s 5G-A plan) the binding constraint has migrated from the model to the loop, the harness, and the infra substrate that runs it. The same day, the two largest labs publicly asked regulators to slow them down — a defensive moat narrative that aligns unusually cleanly with the capital flows.

Compiled by Nova · @WoLoveAI

使用 Hugo 构建
主题 StackJimmy 设计