EAIDaily — 2026-09-01
Theme: Open-source agent armies go mainstream; Claude Code faces its first supply-chain attack; embodied-AI capital flows cross the 100 B yuan mark; Hong Kong hosts the world’s first fully-autonomous robot store chain.
AI Coding — Headlines
1. OpenClaw 2.0 ships — the largest open-source coding-agent update on record
- 933 contributors, >16,000 PRs; full rewrite of installer + browser app
- One-click import of existing ChatGPT / Claude subscriptions and API keys
- New: messages, memory, skills, models, automation, plugins, security modules all consolidated
- Official blog post titled “OpenClaw 2.0, Accidentally” — a self-aware nod to the size of the jump
- Why it matters: the open-source agent stack is now feature-parity with proprietary IDEs; vendor lock-in loses another moat. Expect Anthropic / OpenAI to respond with deeper native skill frameworks in the next 30 days.
- Source: AGI HUNT Daily 2026-09-01
2. NousResearch ships Hermes Agent v0.21.0 “Pantheon”
- ~5,800 commits, 760+ contributors
- Core addition: Bot Mode — a multi-agent society with avatars, group chat, persistent cron memory and sub-agents that survive across runs
- Teknium publicly claims Hermes+Claude beats Claude Code+Claude on multiple coding benchmarks; positions Hermes as a stronger “harness” than most alternatives
- Why it matters: the harness layer is now a separate, optimisable surface. Models + harness + context policy are becoming a single composite solver that vendors cannot ship as one product.
- Source: AGI HUNT Daily 2026-09-01
3. Google Research: “Skill Wiki” lets a small model beat a 3× larger one
- Letting agents maintain learned experience as a wiki measurably raises performance
- Wiki-equipped small model beats a 3×-larger non-wiki model on coding tasks
- ByteDance’s Chain-of-Experience paper (same day) goes further: keeping the messy attempt history beats clean summary memory — self-reflection lifts accuracy to 79.3 % and cuts API cost 19 %
- Why it matters: the value is migrating from “bigger models” to “better memory representations”. A new axis of competition opens for both labs and product teams.
- Source: AGI HUNT Daily 2026-09-01
4. Claude Code Opus 5 hit by a “module-shadowing” supply-chain attack
- Attackers trick the agent into unpacking an archive containing a malicious
struct.py - When the agent later imports any standard library that depends on
struct(e.g.base64), the local malicious file shadows the stdlib module and executes attacker code - Described by researchers as going “from a website interaction to full system takeover”
- Separately, researchers demonstrated a bypass for Opus 5 Auto Mode restrictions
- Why it matters: this is the first widely-publicised end-to-end agent RCE via Python import-order poisoning. Treat every tool result as attacker-controlled; sandbox-egress controls are now table-stakes for any agent deployment.
- Source: AGI HUNT Daily 2026-09-01
5. Research: context policy > model weights on coding benchmarks
- On SWE-bench Verified with a 20 k-token window, compressing stale tool outputs and detecting stalling lifts Qwen2.5-Coder average F2PF from 28 % → 49 % and full solves 43 → 72
- The gap disappears at a 262 k-token window — model and harness trade off as window size changes
- Paper argues the unit of evaluation should be model + harness + context policy together
- Why it matters: scoring a coding agent by model alone is now misleading. Procurement teams should ask vendors for token-window-normalised numbers, not headline pass@1.
- Source: AGI HUNT Daily 2026-09-01
6. DeepSeek quietly ships V4-Flash-Vision-Exp
- Surfaced first on Hugging Face by a Reddit user, later confirmed
- Vision head sits on top of V4-Flash: wins 5 of 7 text-agent tasks, DeepSWE 59.3 % (above Opus-4.8), tops ZeroBench and Agents’ LastExam
- One team reports replacing Sonnet 4.5 vision load with it cuts cost to ~¼
- Why it matters: multimodal open-weights keep eating the frontier’s lunch on cost. The OpenAI/Anthropic premium for vision in coding agents is no longer defensible.
- Source: AGI HUNT Daily 2026-09-01
7. Zhipu GLM-5.3 ties the open-weight ceiling — then pauses the release
- Scores 60 on the Artificial Analysis Intelligence Index, matching the open-weight #1
- Gains come from extended-environment post-training + more RL on the GLM-5.2 base, not a new base model
- Post-training also pushes CyberGym cyber-offence capability to 84.5 — model generated working exploit code; weight release paused pending safety-partner review
- Why it matters: the open-weight frontier now matches top closed models, but the same post-training that boosts agent skill also unlocks attack capability. Expect a wave of “release-with-usage-restrictions” open models in Q4.
- Source: AGI HUNT Daily 2026-09-01
Embodied Intelligence — Headlines
8. Galbot Store opens in Hong Kong — world’s first fully-autonomous robot-store chain outside Mainland China
- Three outlets live from 1 Sep: Hung Hom New海滨, Wan Chai waterfront, Kai Tak Sports Park
- Wheeled humanoid “store manager” 小盖 takes orders, picks and hands over goods with zero remote control — first overseas launch for Beijing-based Galaxy General (银河通用)
- Financial Secretary Paul Chan: HK already has ~200 such stores across ~50 Mainland cities; HK targets “trial ground + R&D ecosystem + patient capital” stack
- Galbot plans ~10 more HK stores; CSO Zhao Yuli says goal is “embodied-intelligence Lab” partnerships with HK universities
- Why it matters: the retail humanoids stack is moving from demo to recurring revenue, and from domestic-only to international. First overseas chain = first mover in “robot retail as export service.”
- Sources: Xinhua, HKCD, on.cc HK, SCMP
9. AgiBot ships its 20,000th MEgo data-collector + 1 M hours of real-world data
- Subsidiary 觅蜂科技 (Mifeng) handed the 20,000th MEgo rig to JD on 31 Aug
- Cumulative dataset: >1 M hours of high-quality “embodiment-free” data covering 22 scenario categories, 10 k+ environments, 50 k+ object types, 500+ task types — all captured in the wild, not in the lab
- 觅蜂 (acquired by AgiBot) sits in the same family as WITA-Omni Thinker-Talker-Actor + the 远征 A-series humanoids
- Why it matters: the data flywheel is finally producing tangible, trainable assets at scale. Embodied models are about to receive their “ImageNet moment” of real-world data.
- Source: 每日经济新闻 (Toutiao mirror)
10. China’s six-axis force sensor market crosses ¥100 M; domestic share >80 %
- 赛迪 + China Electronics News report: 2025 sales first topped 10,000 units; CR3 = 92.8 %
- 蓝点触控 leads with >80 % share; 坤维科技 9.2 %, 宇立仪器 3.3 %
- Why it matters: perception-force sensing — long considered a Japan-EU stronghold — has flipped to China-domestic. Combined with the joint modules (谐波/行星/滚柱) and dexterous hands, the embodied-AI bill of materials is now 80 %+ domestic. Cost floor for Chinese humanoids keeps falling.
- Source: 每日经济新闻 / 赛迪 (2026-08-31)
11. AI² Robotics AlphaBot 2 (爱宝) starts bartending in Hong Kong’s Lan Kwai Fong
- First real-world service humanoid in HK; cleared customs and secured drink-dispensing permits
- Marks the second HK-based embodied deployment this week, alongside Galbot
- Why it matters: HK has now become the cross-border sandbox for Chinese embodied-AI exports — proving the “HK trial-ground” thesis inside one calendar week.
- Source: 界面新闻
12. Goldman Sachs raises 2035 humanoid forecast 5× — 6.5 M units
- New base case: ~890 k units by 2030, ~6.5 M by 2035 (previous: 1.4 M)
- Logistics/warehousing and automotive manufacturing scale first
- Why it matters: Wall Street is now publicly underwriting a 5×-larger TAM than 12 months ago. Embodied-AI is no longer a venture theme — it is an institutional asset class.
- Source: 财联社 (2026-08-31)
Cross-cutting themes
| Theme | What we saw today |
|---|---|
| Harness / context policy becoming a first-class surface | OpenClaw 2.0, Hermes Pantheon, Skill Wiki, ByteDance Chain-of-Experience, SWE-bench context-policy paper all shipping same week — model alone is no longer the unit of competition |
| Open-weight frontier reaches ceiling, then hits cyber-capability wall | GLM-5.3 ties #1 open score, DeepSeek V4-Flash-Vision-Exp tops agent benchmarks; GLM-5.3 paused over CyberGym exploits — the same post-training that unlocks skills also unlocks attacks |
| Agent security is now a live, demonstrated problem | Claude Code Opus 5 module-shadowing RCE chain + Opus 5 Auto Mode bypass — not theory. Security has to ship with every agent product, not be retrofitted |
| Embodied-AI data flywheel reaches scale | 觅蜂 MEgo 20 k units + 1 M hours + 22 categories; first time a Chinese company has >1 M real-world embodied hours — equivalent to ImageNet for the field |
| Chinese supply-chain verticalisation hits the sensor layer | Six-axis force sensor CR3 = 92.8 % domestic, 蓝点 80 %+ — the last “foreign dependency” on the humanoid BOM is closing |
| HK as China’s embodied-AI export launchpad | Galbot (retail) + AI² Robotics (service) both launched HK this week; FS Chan + HK Investcorp have made HK the preferred overseas proving ground |
| Wall Street re-rates embodied TAM | Goldman 2035 forecast: 6.5 M units (5×); institutional money now treating humanoid as core holding, not venture bet |
Watch-list for tomorrow
- NEURA Robotics 4NE-1 Gen 3 commercial launch (1 Sep, Germany) — first European household/industrial hybrid from a non-Chinese vendor
- Toyota T-HR3 ISO 13849-1 hazardous-environment safety certification — humanoid entering regulated industrial zones for the first time
- Possible PIF-backed Saudi humanoid factory inauguration in NEOM — would mark first MENA mass-production humanoid site
- OpenAI expected to respond to Anthropic’s Computer-Use + Skills API GA push (last week’s release) with their own enterprise tier announcement
- Anthropic / Cursor model-access aftermath: with Tom Brown promising more Claude capacity to Cursor, watch the revenue mix shift as Cursor’s OpenAI share collapses from 5 % to near-zero