今日重点项目雷达
DietrichGebert/ponytail
是什么:Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.;判断标签:趋势增强
Source ↗benchflow-ai/awesome-evals
是什么:A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.;判断标签:新变量
Source ↗amplifthq/opentag
是什么:Open-source @agent mentions for Slack and GitHub. OpenTag routes tagged requests to Codex, Claude Code, then returns results in thread.;判断标签:新变量
Source ↗helloianneo/ian-xiaohei-illustrations
是什么:中文小黑怪诞正文配图生成 Skill | 16:9 白底手绘 | 少量红橙蓝批注 | Codex Skill;判断标签:趋势增强
Source ↗omnigent-ai/omnigent
是什么:Omnigent is an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents — swap harnesses without rewriting, enforce policies and sandboxing, and collaborate in real time from any device.;判断标签:趋势增强
Source ↗nexu-io/html-video
是什么:Programmatic video for coding agents — HTML to video on your laptop. Turn HTML, CSS & data into real MP4s with pluggable render engines, 21 templates, AI soundtrack. Apache-2.0, no per-render fees. An official project by the Open Design team.;判断标签:趋势增强
Source ↗NO6KIKO/gorest-2d-animation-spritesheet-generator
是什么:Codex-assisted local 2D animation spritesheet generator and scene compositing workspace.;判断标签:待验证
Source ↗QwenLM/Qwen-AgentWorld
是什么:Qwen-AgentWorld: Language World Models for General Agents;判断标签:趋势增强
Source ↗orange2ai/renwei-writing
是什么:人味儿写作 · An AI agent skill: edit people's words without erasing the person behind them;判断标签:新变量
Source ↗数据来源:AI HOT + GitHub 近期爆发项目 + BestBlogs 今日精选 + Follow Builders 建造者 feed + 必要核实。筛选标准:实时性、增速、AI/Agent 相关度、对 ABU9/OpenClaw 的启发。采集时间:2026-06-27 04:00 Asia/Shanghai;采集窗口:过去 24 小时与 GitHub 7/14/30/90 天爆发窗口。
一、今日主线判断
- 前沿模型进入“发布受控”叙事。 OpenAI 官方确认 GPT-5.6 Sol/Terra/Luna 先面向少量 trusted partners 做有限预览,且该节奏来自与美国政府的发布前沟通和请求。企业不能只等一个闭源最强模型。
- Agent 工程化从工具调用转向运行时治理。 OpenRouter MCP、Claude Code 权限/OTel 更新、OpenTag 都在补模型路由、权限解释、线程回写和可观测性。
- Skill 正在变成轻量产品形态。 配图、写作、PPT、动画、投研、教程类 skill 连续爆发,短期商业价值会先出现在“可交付资产”而非通用聊天。
- 汽车 AI 的监管与 VLA/VLM 叙事同时升温。 小鹏与 WP29/UNR 规则相关消息提示,车端智能与经销商/售后数字化会更快汇合。
- 评测和来源治理的重要性继续上升。 Agent eval、AI 政治倾向测试、版权诉讼共同提醒:没有数据来源、质量门和可信度标签的 AI 输出不能进生产。
二、GitHub 近期爆发项目
DietrichGebert/ponytail
- 是什么:Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
- 元数据:创建 2026-06-12;stars 59,924;最近更新 2026-06-26;采集窗口/来源查询 GitHub Search:created:>=2026-03-29 pushed:>=2026-06-20 stars:>500 AI/agent/LLM/Codex/Claude;age_days 14;repeat_count_7d 2。
- 判断标签:趋势增强
- 意味着什么:“少写代码”的 agent skill 叙事继续放大,说明 Coding Agent 的下一层竞争不只是模型能力,而是可复用工作哲学和约束。OpenClaw 可把它当作 skill 治理样例观察。
- 链接:https://github.com/DietrichGebert/ponytail
benchflow-ai/awesome-evals
- 是什么:A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.
- 元数据:创建 2026-06-24;stars 451;最近更新 2026-06-26;采集窗口/来源查询 GitHub Search:created:>=2026-06-20 stars:>100;age_days 2;repeat_count_7d 1。
- 判断标签:新变量
- 意味着什么:Agent eval 资源在新仓库上快速聚合,说明企业开始从“能跑”转向“能评估、能复盘、能选型”。ABU9 的 AI 交付需要同步建立评测清单。
- 链接:https://github.com/benchflow-ai/awesome-evals
amplifthq/opentag
- 是什么:Open-source @agent mentions for Slack and GitHub. OpenTag routes tagged requests to Codex, Claude Code, then returns results in thread.
- 元数据:创建 2026-06-24;stars 286;最近更新 2026-06-26;采集窗口/来源查询 GitHub Search:created:>=2026-06-20 stars:>100;age_days 2;repeat_count_7d 0。
- 判断标签:新变量
- 意味着什么:Slack/GitHub 中直接 @agent 路由到 Codex、Claude Code 的模式很接近企业协同入口,值得观察其权限、线程回写和审计方式。
- 链接:https://github.com/amplifthq/opentag
helloianneo/ian-xiaohei-illustrations
- 是什么:中文小黑怪诞正文配图生成 Skill | 16:9 白底手绘 | 少量红橙蓝批注 | Codex Skill
- 元数据:创建 2026-05-27;stars 6,266;最近更新 2026-06-26;采集窗口/来源查询 GitHub Search:created:>=2026-05-28 stars:>500 AI/agent/LLM/Codex/Claude;age_days 29;repeat_count_7d 0。
- 判断标签:趋势增强
- 意味着什么:中文内容配图 skill 高速传播,说明“文章到可发布视觉资产”的链路正在变成轻量产品。对 ABU9 市场材料和客户汇报有直接可复制价值。
- 链接:https://github.com/helloianneo/ian-xiaohei-illustrations
omnigent-ai/omnigent
- 是什么:Omnigent is an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents — swap harnesses without rewriting, enforce policies and sandboxing, and collaborate in real time from any device.
- 元数据:创建 2026-06-11;stars 4,990;最近更新 2026-06-26;采集窗口/来源查询 GitHub Search:created:>=2026-05-28 stars:>500 AI/agent/LLM/Codex/Claude;age_days 15;repeat_count_7d 3。
- 判断标签:趋势增强
- 意味着什么:多 harness / 多 agent 编排仍有热度,但 repeat_count 较高,今天的新信息主要是持续活跃与 star 增速;需要验证真实运行质量。
- 链接:https://github.com/omnigent-ai/omnigent
krea-ai/krea-2
- 是什么:Official inference code for Krea 2
- 元数据:创建 2026-06-22;stars 303;最近更新 2026-06-26;采集窗口/来源查询 GitHub Search:created:>=2026-06-20 stars:>100;age_days 3;repeat_count_7d 0。
- 判断标签:新变量
- 意味着什么:图像/视频模型官方推理代码开源,表示创意模型正在用代码仓库建立开发者入口;可关注其是否进入企业素材生成流水线。
- 链接:https://github.com/krea-ai/krea-2
nexu-io/html-video
- 是什么:Programmatic video for coding agents — HTML to video on your laptop. Turn HTML, CSS & data into real MP4s with pluggable render engines, 21 templates, AI soundtrack. Apache-2.0, no per-render fees. An official project by the Open Design team.
- 元数据:创建 2026-05-27;stars 3,562;最近更新 2026-06-26;采集窗口/来源查询 GitHub Search:created:>=2026-05-28 stars:>500 AI/agent/LLM/Codex/Claude;age_days 30;repeat_count_7d 0。
- 判断标签:趋势增强
- 意味着什么:HTML 到 MP4 的 agent-friendly 生产链路与日报、销售视频、培训短片相关,是 OpenClaw artifact pipeline 可吸收的方向。
- 链接:https://github.com/nexu-io/html-video
NO6KIKO/gorest-2d-animation-spritesheet-generator
- 是什么:Codex-assisted local 2D animation spritesheet generator and scene compositing workspace.
- 元数据:创建 2026-06-12;stars 984;最近更新 2026-06-26;采集窗口/来源查询 GitHub Search:created:>=2026-06-13 stars:>200 AI/agent/LLM/Codex/Claude;age_days 14;repeat_count_7d 0。
- 判断标签:待验证
- 意味着什么:Codex-assisted 动画资产生成说明游戏/营销素材制作正在被 skill 化,但需要验证输出质量和版权边界。
- 链接:https://github.com/NO6KIKO/gorest-2d-animation-spritesheet-generator
QwenLM/Qwen-AgentWorld
- 是什么:Qwen-AgentWorld: Language World Models for General Agents
- 元数据:创建 2026-06-22;stars 561;最近更新 2026-06-26;采集窗口/来源查询 GitHub Search:created:>=2026-06-13 stars:>200 AI/agent/LLM/Codex/Claude;age_days 4;repeat_count_7d 2。
- 判断标签:趋势增强
- 意味着什么:通义系把“语言世界模型 for general agents”做成开源项目,强化中国开源 agent 训练/评测路线。
- 链接:https://github.com/QwenLM/Qwen-AgentWorld
orange2ai/renwei-writing
- 是什么:人味儿写作 · An AI agent skill: edit people's words without erasing the person behind them
- 元数据:创建 2026-06-12;stars 872;最近更新 2026-06-26;采集窗口/来源查询 GitHub Search:created:>=2026-06-13 stars:>200 AI/agent/LLM/Codex/Claude;age_days 13;repeat_count_7d 0。
- 判断标签:新变量
- 意味着什么:写作类 skill 从“润色”转向“保留人味”,对企业内部知识输出、领导讲话稿和客户材料更实际。
- 链接:https://github.com/orange2ai/renwei-writing
三、最近 7 天新项目:值得点开
winsznx/theeleven
- 是什么:Eleven autonomous AI agents open live football prop markets on X Layer — custom Uniswap v4 hook, gasless USDT0 staking.
- 元数据:创建 2026-06-25;stars 459;最近更新 2026-06-26;采集窗口/来源查询 GitHub Search:created:>=2026-06-20 stars:>100;age_days 1;repeat_count_7d 0。
- 判断标签:待验证
- 意味着什么:自主 agent + 链上预测市场叙事很强,但与企业 AI 相关度低,先作为“agent 叙事外溢到金融/娱乐”的观察点。
- 链接:https://github.com/winsznx/theeleven
amplifthq/opentag
- 是什么:Open-source @agent mentions for Slack and GitHub. OpenTag routes tagged requests to Codex, Claude Code, then returns results in thread.
- 元数据:创建 2026-06-24;stars 286;最近更新 2026-06-26;采集窗口/来源查询 GitHub Search:created:>=2026-06-20 stars:>100;age_days 2;repeat_count_7d 0。
- 判断标签:新变量
- 意味着什么:2 天内获得关注,企业协作入口价值高,建议优先看安装体验、权限和线程回写。
- 链接:https://github.com/amplifthq/opentag
benchflow-ai/awesome-evals
- 是什么:A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.
- 元数据:创建 2026-06-24;stars 451;最近更新 2026-06-26;采集窗口/来源查询 GitHub Search:created:>=2026-06-20 stars:>100;age_days 2;repeat_count_7d 1。
- 判断标签:趋势增强
- 意味着什么:新建 2 天即高关注,说明 agent eval 需求强,适合作为 ABU9 AI 交付评测资料入口。
- 链接:https://github.com/benchflow-ai/awesome-evals
QwenLM/Qwen-AgentWorld
- 是什么:Qwen-AgentWorld: Language World Models for General Agents
- 元数据:创建 2026-06-22;stars 561;最近更新 2026-06-26;采集窗口/来源查询 GitHub Search:created:>=2026-06-13 stars:>200 AI/agent/LLM/Codex/Claude;age_days 4;repeat_count_7d 2。
- 判断标签:趋势增强
- 意味着什么:4 天内进入高关注,适合跟踪其 benchmark、训练 recipe 与中文 agent 生态联动。
- 链接:https://github.com/QwenLM/Qwen-AgentWorld
bozhouDev/codex-orange-book
- 是什么:Codex 橙皮书:从安装到实战案例的全链路 Codex 使用指南(非官方开源,含可下载 PDF)
- 元数据:创建 2026-06-23;stars 2,125;最近更新 2026-06-26;采集窗口/来源查询 GitHub Search:created:>=2026-05-28 stars:>500 AI/agent/LLM/Codex/Claude;age_days 3;repeat_count_7d 3。
- 判断标签:趋势增强
- 意味着什么:重复出现已达 3 次,今天降为“持续跟踪”:价值在中文 Codex 实战案例,但需要避免把教程热度误判为底层技术突破。
- 链接:https://github.com/bozhouDev/codex-orange-book
WangJunqing-coder/huasheng13-skill
- 是什么:基于花生十三公开教学体系、课程资料、历年正题整理得到的skill。
- 元数据:创建 2026-06-22;stars 299;最近更新 2026-06-26;采集窗口/来源查询 GitHub Search:created:>=2026-06-20 stars:>100;age_days 4;repeat_count_7d 0。
- 判断标签:待验证
- 意味着什么:知识型 skill 快速涌现,验证点不在 star,而在素材来源、版权、可追溯引用和是否可复用到企业培训。
- 链接:https://github.com/WangJunqing-coder/huasheng13-skill
SarkAzia/baiyueguang-learning-skill
- 是什么:用视奸前任白月光的方式学数学以及任何学科
- 元数据:创建 2026-06-21;stars 290;最近更新 2026-06-26;采集窗口/来源查询 GitHub Search:created:>=2026-06-20 stars:>100;age_days 5;repeat_count_7d 0。
- 判断标签:噪音降权
- 意味着什么:传播标题强但企业价值弱,暂作为 skill 爆款包装方式观察,不进入技术主线。
- 链接:https://github.com/SarkAzia/baiyueguang-learning-skill
BohemiaInteractive/CWR
- 是什么:Arma: Cold War Assault Remastered Source Code Repository.
- 元数据:创建 2026-06-22;stars 603;最近更新 2026-06-26;采集窗口/来源查询 GitHub Search:created:>=2026-06-20 stars:>100;age_days 4;repeat_count_7d 0。
- 判断标签:噪音降权
- 意味着什么:非 AI 仓库,只因 GitHub 新仓热度进入候选;用于提醒采集结果必须人工降噪。
- 链接:https://github.com/BohemiaInteractive/CWR
四、AI 热点新闻
OpenAI 预览 GPT-5.6 系列,并确认有限发布节奏
- 事件:OpenAI 预览 GPT-5.6 Sol/Terra/Luna,称 Sol 是旗舰模型、Terra 面向日常工作、Luna 主打快速低成本;初期仅向少量 trusted partners 和组织开放 API/Codex 访问。
- 可信度:一手发布
- 意味着什么:OpenAI 已披露初步能力、定价和缓存机制,但广泛可用仍未开始;不要把它当作今天即可采购和交付的稳定能力。
- 链接:https://openai.com/index/previewing-gpt-5-6-sol
GPT-5.6 有限预览来自美国政府请求
- 事件:OpenAI 官方称,在与美国政府持续沟通后,应其请求先向少量 trusted partners 做有限预览,后续再扩大发布。
- 可信度:一手发布
- 意味着什么:前沿模型发布正在进入“能力审查/客户准入/分阶段开放”的窗口期。ABU9/OpenClaw 要把多模型路由、开源或国产替代、供应商降级路径写进默认架构。
- 链接:https://openai.com/index/previewing-gpt-5-6-sol
OpenRouter MCP Server 发布
- 事件:OpenRouter MCP Server 发布。
- 可信度:一手发布
- 意味着什么:模型选择、价格、benchmark 与测试推理被放进 IDE/MCP,模型路由正在从平台后台变成 agent 的实时工具。
- 链接:https://openrouter.ai/blog/announcements/openrouter-mcp-server
Codex 在 ChatGPT 移动 App 正式可用
- 事件:Codex 在 ChatGPT 移动 App 正式可用。
- 可信度:一手发布
- 意味着什么:移动端启动、审阅和批准后台任务,会扩大非工程岗位使用 Codex 的场景;OpenClaw 也要考虑移动审批和任务状态通知。
- 链接:https://x.com/OpenAIDevs/status/2070254532911882707
Claude Code v2.1.193 增强 auto mode、权限解释和 OTel 日志
- 事件:Claude Code v2.1.193 增强 auto mode、权限解释和 OTel 日志。
- 可信度:一手发布
- 意味着什么:Claude Code 正在补齐安全分类、可观测性和后台任务稳定性;这是 OpenClaw agent runtime 需要对标的工程方向。
- 链接:https://github.com/anthropics/claude-code/releases/tag/v2.1.193
Runway Agent 2.0 面向营销生产链路
- 事件:Runway Agent 2.0 面向营销生产链路。
- 可信度:一手发布
- 意味着什么:视频/广告/本地化/投放分析被整合为 agent 工作流,说明营销 Agent 的可付费点是资产迭代而不是单次生成。
- 链接:https://runwayml.com/news/introducing-agent-2
Anthropic Economic Index 发布 Claude 使用节奏报告
- 事件:Anthropic Economic Index 发布 Claude 使用节奏报告。
- 可信度:一手发布
- 意味着什么:隐私保护遥测正在变成 AI 经济研究方法;但自动化乐观预期属于调查结论,仍需长期验证。
- 链接:https://www.anthropic.com/research/economic-index-june-2026-report
小鹏称 2026 年底自动驾驶可合法进入全球
- 事件:小鹏称 2026 年底自动驾驶可合法进入全球。
- 可信度:媒体报道
- 意味着什么:与 WP29/UNR 规则相关,汽车 AI 的监管窗口在打开;ABU9 应关注经销商端 VLA/VLM 体验如何落地到销售和售后。
- 链接:https://www.ithome.com/0/968/894.htm
近 400 家美国报纸起诉微软和 OpenAI
- 事件:近 400 家美国报纸起诉微软和 OpenAI。
- 可信度:媒体报道
- 意味着什么:版权诉讼继续抬高企业 AI 内容合规成本;知识库/训练/引用链路要保留来源和授权边界。
- 链接:https://www.ithome.com/0/968/872.htm
General Intuition 获 3.2 亿美元融资,用游戏数据训练通用 agent
- 事件:General Intuition 获 3.2 亿美元融资,用游戏数据训练通用 agent。
- 可信度:媒体报道
- 意味着什么:融资与估值不能等同技术验证;真正值得看的是“带动作标签的视频数据”是否能迁移到机器人和复杂操作。
- 链接:https://techcrunch.com/2026/06/25/from-fortnite-to-robots-general-intuitions-2-3b-bet-that-video-games-can-train-ai-agents-for-the-real-world
五、论文 / 研究 / 观点
Meta 隐私感知基础设施资产分类:LLM 先处理歧义,再蒸馏为确定性规则
- 判断:对企业 AI 的启发很强:不要让 LLM 直接做生产决策,先让它帮助生成规则、证据和评估集。
- 链接:https://engineering.fb.com/2026/06/25/security/privacy-aware-infrastructure-in-the-ai-native-era-an-asset-classification-case-study
GitHub Copilot agentic harness 跨模型/任务评估
- 判断:模型不是唯一变量,harness 的 token 效率、工具设计和任务分解同样决定成本。
- 链接:https://github.blog/ai-and-ml/github-copilot/evaluating-performance-and-efficiency-of-the-github-copilot-agentic-harness-across-models-and-tasks
OLMo Hybrid vs Transformer:混合架构在实义词上优势更明显
- 判断:架构路线仍在分化,企业选型应看任务类型,不要把单一 benchmark 当成通用结论。
- 链接:https://huggingface.co/blog/allenai/hybrid-token-prediction
AI 聊天机器人政治倾向测试引发讨论
- 判断:这类测试对“默认回答框架”有参考价值,但政治立场属于高敏感评测,结论应作为风险观察而非产品排名。
- 链接:https://the-decoder.com/most-major-ai-chatbots-still-lean-left-on-political-questions-even-anti-woke-models-are-no-exception
Tomer Tunguz:异步、多轮 agent 推理需要 fleet-aware 编排
- 判断:长任务 agent 的成本瓶颈会从单次 token 价格转到队列、吞吐、失败恢复和调度。
- 链接:https://www.tomtunguz.com/sail-inference-queue
六、今天值得读 / 值得看(BestBlogs / Follow Builders 精选)
标题/核心观点:Peter Yang:让 Codex 用浏览器规划旅行并保存价格/链接
- 为什么值得看:浏览器 Agent 已经从 demo 走向个人可用工作流,ABU9 可映射到竞品巡检、经销商网页检查、售后流程核验。
- 原帖链接:https://x.com/petergyang/status/2070353698140958818
标题/核心观点:Aaron Levie:前沿模型发布进入事实监管时代
- 为什么值得看:值得看它对闭源模型、开源权重、主权 AI 和企业采购风险的结构化拆解。
- 原帖链接:https://x.com/levie/status/2070310706369712272
标题/核心观点:Guillermo Rauch:Next.js 错误修复提示加入 Copy prompt
- 为什么值得看:软件产品开始主动为 agent 生成上下文,企业系统也应在错误页、日志页、流程页提供 agent-readable 信息。
- 原帖链接:https://x.com/rauchg/status/2070243120546218000
标题/核心观点:Guillermo Rauch:把设计标准注入 coding agents
- 为什么值得看:对 OpenClaw 很直接:设计规范、交付格式和质量门要变成 agent 可执行上下文。
- 原帖链接:https://x.com/rauchg/status/2070241572416078161
标题/核心观点:Claude Code Hook 玩法:事件驱动自动化
- 为什么值得看:Hook 不耗 token,适合把提醒、总结、文件整理、推送变成 agent runtime 的规则层。
- 链接:https://mp.weixin.qq.com/s/LVj2foSXi_hBRKxjuYaUyw
标题/核心观点:Leaf 实时通话 AI 分身工程拆解
- 为什么值得看:价值不在“复刻网红”,而在 ASR/LLM/TTS/persona 的低延迟链路,对智能销售助理可借鉴。
- 原帖链接:https://x.com/AYi_AInotes/status/2070531964067623381
标题/核心观点:小互 IP Studio:文章到配图 skill
- 为什么值得看:内容资产生产的 skill 化很明显,适合 ABU9 做客户案例图、活动海报、公众号配图。
- 原帖链接:https://x.com/xiaohu/status/2070317717811540149
标题/核心观点:Meta PAI 案例:稳定行为规则化
- 为什么值得看:比“让模型更聪明”更重要的是把可稳定复现的行为沉淀为规则、测试和审计。
- 链接:https://engineering.fb.com/2026/06/25/security/privacy-aware-infrastructure-in-the-ai-native-era-an-asset-classification-case-study
七、对我们有用的行动建议
- OpenClaw:把模型路由做成 MCP/工具层能力,记录模型、价格、benchmark、失败率和任务类型,避免人工凭感觉选模型。
- OpenClaw:把 Hook、权限解释、OTel 日志、后台任务回收列为 agent runtime 标配,不要只做聊天入口。
- ABU9:优先做“汽车行业 skill 包”样板:售前方案、经销商活动物料、交付验收清单、车型知识问答、售后工单摘要。
- ABU9:把智能销售助理从文本问答升级为低延迟语音链路验证,重点测首字响应、打断、转人工、知识溯源。
- 企业 AI:所有融资、benchmark、模型能力、政治倾向、异常 star 项目都进入观察层,决策层只接收一手发布或可复核证据。
八、今日结论
今天的结构性判断:AI 的竞争重心正在从“谁的模型更强”转向“谁能把模型放进可审计、可路由、可复用、可交付的运行系统”;OpenClaw 和 ABU9 应押注 agent runtime + 行业 skill + artifact 交付。
降级说明
- BestBlogs:采集接口返回 success=true 但 data=null,本期未能使用 BestBlogs 原始条目;“值得读/值得看”改用 Follow Builders、AI HOT 与一手博客补足。
- GitHub API:采集成功;但部分新仓 star 异常偏高或与 AI 弱相关,已标注“待验证/噪音降权”。
- 新闻可信度:融资、估值、政府审批、未公开模型能力、benchmark 均未与正式发布同权重处理。