每日简报

2026-08-23

← 历史归档

openai/codex

Rust · ★ 113,215 · 🍴 17,350 · 📈 1,978 stars today

Lightweight coding agent that runs in your terminal

中文介绍 轻量级编码代理,运行于终端,可执行代码任务,解释复杂代码,优化git工作流。

mattpocock/skills

Shell · ★ 231,903 · 🍴 19,796 · 📈 2,684 stars today

Skills for Real Engineers. Straight from my .agents directory.

中文介绍 为真实工程师设计的技能集,源自作者的个人技能目录。

affaan-m/ECC

JavaScript · ★ 242,150 · 🍴 36,699 · 📈 428 stars today

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

中文介绍 性能优化系统,具备技能、本能、记忆、安全和研究优先的开发,适用于Claude Code等工具。

obra/superpowers

Shell · ★ 276,156 · 🍴 24,695 · 📈 592 stars today

An agentic skills framework & software development methodology that works.

中文介绍 一个代理技能框架和软件开发方法论,旨在提高开发效率。

Wei-Shaw/sub2api

Go · ★ 38,772 · 🍴 8,037 · 📈 264 stars today

Sub2API 一站式开源中转服务,让 Claude、Openai 、Gemini、Grok订阅统一接入,支持拼车共享,更高效分摊成本,原生工具无缝使用。

中文介绍 一站式开源中转服务,支持Claude、Openai等订阅服务,实现成本分摊和工具无缝使用。

makeplane/plane

TypeScript · ★ 57,196 · 🍴 5,439 · 📈 263 stars today

🔥🔥🔥 Open-source Jira, Linear, Monday, and ClickUp alternative. Plane is a modern project management platform to manage tasks, sprints, docs, and triage.

中文介绍 开源的Jira、Linear、Monday和ClickUp替代品,现代项目管理平台,支持任务、迭代、文档和问题排序。

n8n-io/n8n

TypeScript · ★ 201,788 · 🍴 60,298 · 📈 202 stars today

Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.

中文介绍 具有原生AI能力的公平代码工作流自动化平台,结合可视化构建和自定义代码,支持自托管或云部署,提供400+集成。

anthropics/claude-code

Python · ★ 142,505 · 🍴 22,838 · 📈 141 stars today

Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflows - all through natural language commands.

中文介绍 终端中的代理编码工具,理解代码库,通过执行常规任务、解释复杂代码和优化git工作流来加速编码。

AprilNEA/OpenLogi

Rust · ★ 13,853 · 🍴 370 · 📈 959 stars today

⚡️A native, local-first alternative to Logitech Options+, written in Rust 🦀 — remap buttons, DPI, and SmartShift over HID++. No account, no telemetry.

中文介绍 基于Rust的本地化Logitech Options+替代品,支持按钮重映射、DPI和SmartShift,无需账户和遥测数据。

modular/modular

Mojo · ★ 28,828 · 🍴 3,065 · 📈 395 stars today

The Modular Platform (includes MAX & Mojo)

中文介绍 Modular平台,包括MAX和Mojo,提供音乐制作和表演的软件和硬件解决方案。

multica-ai/andrej-karpathy-skills

★ 205,268 · 🍴 21,005 · 📈 379 stars today

A single CLAUDE.md file to improve Claude Code behavior, derived from Andrej Karpathy's observations on LLM coding pitfalls.

中文介绍 基于Andrej Karpathy对LLM编码陷阱的观察,改进Claude Code行为的CLAUDE.md文件。

mahlernim/google-timeline-visualizer

Kotlin · ★ 2,552 · 🍴 299 · 📈 441 stars today

Visualize your year in travel using your Google Location History (Timeline) data

中文介绍 使用Google位置历史记录数据可视化年度旅行。

ripienaar/free-for-dev

HTML · ★ 133,868 · 🍴 14,027 · 📈 915 stars today

A list of SaaS, PaaS and IaaS offerings that have free tiers of interest to devops and infradev

中文介绍 列出对开发人员有价值的免费SaaS、PaaS和IaaS服务。

microsoft/TypeScript

Go · ★ 110,526 · 🍴 13,746 · 📈 163 stars today

TypeScript is a superset of JavaScript that compiles to clean JavaScript output.

中文介绍 TypeScript是JavaScript的超集,编译为干净的JavaScript输出。

cursor/plugins

TypeScript · ★ 4,647 · 🍴 380 · 📈 286 stars today

Cursor plugin specification and official plugins

中文介绍 Cursor插件规范和官方插件。

PostHog/posthog

Python · ★ 38,589 · 🍴 3,244 · 📈 288 stars today

🦔 PostHog is the leading platform for building self-driving products. Our developer tools – AI observability, analytics, session replay, flags, experiments, error tracking, logs, and more – capture all the context agents need to diagnose problems, uncover opportunities, and ship fixes. Steer it all

中文介绍 PostHog是一个领先的自驾驶产品构建平台,提供AI可观察性、分析、会话回放、标志、实验、错误跟踪、日志等功能。

Tencent/AI-Infra-Guard

Python · ★ 5,473 · 🍴 518 · 📈 161 stars today

A full-stack AI Red Teaming platform securing AI ecosystems via Agent Scan, Skills Scan, MCP scan, AI Infra scan and LLM jailbreak evaluation.

中文介绍 全栈AI红队平台,通过代理扫描、技能扫描、MCP扫描、AI基础设施扫描和LLM越狱评估来保护AI生态系统。

FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving

👍 14

Long-context modeling is a pivotal capability for Large Language Models, yet the quadratic complexity of attention remains a critical bottleneck, particularly during the compute-intensive prefilling phase. Our previous work, FlashPrefill, mitigates this cost through instantaneous pattern discovery a

SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?

👍 58

Software increasingly functions as part of the scientific instrument itself, making failures in scientific code capable of compromising not only program behavior but also the evidence underlying scientific conclusions. Yet existing evaluations of coding agents largely emphasize aggregate task succes

Repo0: Design-Driven Zero-to-All Code Generation

👍 17

Large language model agents have made substantial progress in code generation, yet most existing systems assume a predefined repository architecture. This assumption does not hold in zero-to-all code generation, where an agent must construct an entire software project directly from natural-language

MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use

👍 31

Memory has become a key component of large language models, enabling them to retain information and learn from long-term interactions. However, existing memory benchmarks mainly evaluate whether information is correctly extracted, stored, and retrieved, while largely overlooking how retrieved memori

EnvHarness: Awakening Static Worlds for Agent Learning

👍 246

LLM agents learn by interacting with environments, yet these environments are hand-built and static: blind to an agent's weaknesses, and quickly left behind as it improves. While recent environment generation methods attempt to address this, they require domain-specific pipelines, rely on expensive

FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis

👍 112

Training terminal agents requires scalable executable supervision, yet synthesizing high-quality terminal tasks remains challenging. Each task couples an instruction, an initialized environment, a reference solution, and an executable verifier; if these artifacts are generated from inconsistent assu

SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents

👍 6

Agent frameworks increasingly package procedural knowledge as skills: instruction files an agent reads on demand, while public libraries now hold thousands of them. Which skill to read has thus become a decision the policy itself makes in the middle of an episode, yet no existing signal trains it. W

Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL

👍 91

Reinforcement learning (RL) has emerged as a powerful approach for improving reasoning in language and vision-language models, yet its strongest successes still depend heavily on ground-truth supervision (e.g., verifiable reward). Such annotations are costly to obtain and become increasingly scarce

Chain-of-Experience for Continual LLM Improvement

👍 7

Humans continuously learn from experience, whereas conventional large language model (LLM) evaluations ignore the models' ability to improve through inference-time interaction. In this paper, we study how LLMs learn from iterative experience at test time, a setting we refer to as Chain-of-Experience

Towards Real-Time and Adaptable LiDAR Scene Completion

👍 3

LiDAR scene completion is a key component of 3D perception in autonomous driving, where the scene must be completed in real time to be usable in downstream tasks. Existing approaches typically follow an initialize-and-refine paradigm, in which a coarse initialization of the scene is first constructe

LLMs Get Smarter from Targeted Synthetic Multilingual Data

👍 4

Language-specific competency (LSC) is the phenomenon of a language model performing better or worse depending on the language of the prompt. In other words, a language model outputs different (and potentially incorrect) responses to the same semantic query when prompted in different languages. Prior

Bounded Agents: Delegation Security for Multi-Agent AI Systems

👍 3

LLM-based agents can act on behalf of a user to access cloud services, call tools, or invoke agents. At session start, the agent's permissions are set but remain static, and each request is evaluated independently, without considering prior actions. Within its permissions, an agent may act contrary

The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning

👍 17

Popular facts are memorised more deeply during pretraining and resist removal longer than rare ones, yet existing LLM unlearning methods apply uniform gradient pressure regardless of training-data frequency. We propose the AdaPop (Adaptive Popularity) method, which combines local token confidence wi

The Embedder's Dilemma: LLMs Are Better, but at What Cost?

👍 7

Should you replace your text-embedding pipeline with a large language model? We answer this with a controlled, cost-aware comparison of ten LLMs across six families and 26 embedding models (118M to 14B parameters) on 37 tasks spanning classification, semantic textual similarity (STS), clustering, pa

QuoteBench: How Matched Scores Can Hide Command-Path Failures

👍 7

LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched execution scores alone cannot distinguish command-generation errors from failures introduced after generation. QuoteBench measures this boundary with exact final-state validation on 5

SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback

👍 29

Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently possess no closed loop through which they might improve from the interaction failures they actually cause. Recent work does close this loop, but derives its feedback from single-turn question-an

OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

👍 88

Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and code execution to manuscript preparation. Yet workflow coverage alone does not provide access to the full evidence on which scientific discovery depend

Temporal Multi-Signal Fusion for Token-Level Hallucination Detection

👍 3

Token-level hallucination detectors score each token independently from a single signal, and fail exactly when the generating model is confidently wrong. This paper instead treats hallucination as a temporally extended span and detects it by sequence labeling: each token is scored from a 33-dimensio

Update on our post-training effort

@gabepereyra · 11.4K 粉丝 · 616.8K 阅 · 507 赞 · 99 转

Authors: @nikogrupen @ItsJulioPereyra @calvincongelado @vtrengarajan @gabepereyra Over the past six months, Harvey’s research agenda has focused on two goals: Building frontier legal intelligence

中文介绍 探讨 Harvey 团队过去六个月在法律智能领域的两项研究目标:构建前沿法律智能。

How To Set Up Grokbot The Right Way So It Runs Your Business While You Sleep In 2026

@heynavtoor · 149.3K 粉丝 · 161.9K 阅 · 550 赞 · 93 转

Elon Musk just told his followers to try Grokbot after a power user showed it replaced his $10,000 Mac Mini local AI setup (Elon on X, Aug 16). Gavin Baker says his AI usage went up about 100x after

中文介绍 介绍如何正确设置 Grokbot,实现自动化运营,提升 AI 使用效率。

Building software factories (with no slop)

@dzhng · 13.3K 粉丝 · 137.6K 阅 · 502 赞 · 48 转

The amount of code being written today far outpaces our ability to review it. That part isn't controversial. What's more interesting is that we're visibly losing — the amount of AI slop in modern

中文介绍 分析现代软件编写中代码审查能力不足,以及 AI slop 问题。

Agentic coding is now multiplayer: Introducing Slack Code

@SlackHQ · 434.3K 粉丝 · 86.9K 阅 · 512 赞 · 44 转

AI coding is now a team sport. Code channels bring software development out of private tabs, so teams and agents can write, review, and ship code together, in the open. Slack is where agents become

中文介绍 Slack 推出 Slack Code,实现多人协作的 AI 编码。

AI Engineering Skills Map: Building and Deploying AI Applications

@AndrewYNg · 1.8M 粉丝 · 62.4K 阅 · 1.2K 赞 · 189 转

I previously wrote about our AI Engineering Skills Map, with the highest level skills being (i) Building and deploying AI applications, (ii) Software engineering fundamentals, (iii) Using coding

中文介绍 分享 AI 工程师技能图谱,涵盖构建和部署 AI 应用等关键技能。

RL Policy Churn

@ID_AA_Carmack · 3.7M 粉丝 · 58.8K 阅 · 577 赞 · 40 转

I have a good theory about this now. The Phenomenon of Policy Churn https://arxiv.org/pdf/2206.00730 The classic exploration method in RL is “epsilon greedy”, where the agent normally takes the “greedy” action

中文介绍 探讨强化学习中的策略漂移现象及其经典探索方法 epsilon greedy。

Grok Bot vs. Hermes: They Look Similar. Underneath, They’re Making Very Different Bets.

@themahis · 4.7K 粉丝 · 43.0K 阅 · 515 赞 · 37 转

One is trying to make AI teammates feel like a managed product. The other is turning the agent itself into an open, composable runtime. Research note — August 20, 2026: For this comparison, I

中文介绍 比较 Grok Bot 和 Hermes 的不同策略,探讨 AI 同事产品的未来方向。

5 design patterns for long-horizon agent harness

@GoogleCloudTech · 1.3M 粉丝 · 42.9K 阅 · 506 赞 · 87 转

Any agent can look brilliant on a one-shot task. Give it a week of real work and it starts to fall apart. "Long horizon" means the agent keeps going across days and dozens of sessions instead of

中文介绍 提出五个设计模式,以增强长期目标代理的利用。

System Design for Agent Systems (Part 1)

@kmeanskaran · 13.9K 粉丝 · 32.9K 阅 · 518 赞 · 66 转

Put a frontier model into a badly designed agent system and you get a more articulate failure. Nearly all of the engineering sits outside the model, in the scaffolding that has come to be called the

中文介绍 探讨为代理系统进行系统设计的重要性。

The Evolution of the Agent Harness

Models keep absorbing the harness into their weights — soon, it will be a harness for human attention rather than for the model.

中文介绍 模型不断吸收注意力机制,未来将更专注于人类注意力而非模型本身。

Simulation: the new Scaling Law — Joon Sung Park, Simile AI

Simile’s CEO about his journey from the viral Generative Agents to creating 8 Billion Digital Twins of every living human... and why it’s gone from fun exploration to very serious business.

中文介绍 Simile AI的CEO讲述其从病毒式生成代理到创建80亿个数字孪生的历程。

not much happened today

**Ox Alpha** emerged as a mystery model with strong coding and agentic performance, likely a **Zhipu/GLM-family** model such as **GLM-5.3 Vision**. Analysts suggest its gains come from post-training and infrastructure improvements rather than sheer size, based on the **743B base** of **GLM-5.2** wit

中文介绍 Ox Alpha成为一款神秘模型,具有强大的编码和代理性能,可能是一款Zhipu/GLM-family模型。

Debates over AI consciousness are a trap

“Runaway” AI, “rogue” agents, and “autonomous” actors—the current rhetoric would have you believe that AI agents are not only awake and aware, but angry at their creators. Prominent tech leaders such as Demis Hassabis, Dario Amodei, and Sam Altman push for regulation of these seemingly “superhuman”

中文介绍 关于AI意识的辩论是一个陷阱,一些知名科技领导者呼吁对这些“觉醒”的AI代理进行监管。

Unlocking hidden revenue streams with market models

Each day, an airline transports tens of thousands of passengers on hundreds of flights. Often these are not straightforward point-to-point routes, with passengers requiring multiple connections. The airline can consider potentially hundreds of variables to price each of these journeys: demand, seaso

中文介绍 通过市场模型解锁隐藏的收益流。

Introducing AI Futures

Introducing AI Futures, a new OpenAI blog exploring how transformative AI could reshape power, governance, the economy, and individual freedom.

中文介绍 OpenAI推出AI未来博客,探讨AI如何重塑权力、治理、经济和个人自由。

not much happened today

**OpenAI** and **Anthropic** expanded their agent platforms with new desktop features, collaborative editing, and composable APIs like Skills and Files API. **OpenAI** rolled out memory and workflow features in the EEA, UK, and Switzerland. **AT&T** revealed that 40% of employee AI usage routes to o

中文介绍 OpenAI和Anthropic扩展了他们的代理平台,增加了新的桌面功能、协作编辑和可组合API等。

Real Madrid beat Espanyol 2-1 in Jose Mourinho’s first game on return

Carlos Espi scores late to snatch a win for Real Madrid after Alex Calatrava had levelled Jude Bellingham's opener.

中文摘要 皇家马德里在何塞·穆里尼奥回归首场比赛中以2-1战胜皇家社会,卡洛斯·埃斯皮补时阶段进球逆转。

Open-air cinema brings movie nights back to Khartoum

Families are returning to the movies at an open-air cinema in Khartoum, offering residents a brief escape from war.

中文摘要 开罗露天电影院重燃电影之夜,为居民提供逃离战争的短暂时刻。

Trump’s Top Trade Representative Details Offer That Canada Refused

In an interview, Jamieson Greer, Trump’s trade representative, laid out details of the trade offer that the United States made to Canada before talks crumbled.

中文摘要 特朗普的前贸易代表格雷尔详细介绍了美国向加拿大提出的贸易提议,但谈判破裂。

Israeli army and settlers injure several Palestinians across West Bank

Palestinians face attacks and forced displacement as settlers expand control, backed by military raids and inaction.

中文摘要 以色列军队和定居者在西岸伤害了几名巴勒斯坦人,随着定居者扩大控制,巴勒斯坦人面临袭击和强迫迁移。

US Postal Service shares mail-in ballot restrictions despite court ruling

US President Donald Trump has called for restrictions on mail-in voting as part of bid to exert control over elections.

中文摘要 尽管法院裁决,美国邮政服务仍分享了对邮寄选票的限制,唐纳德·特朗普总统呼吁对邮寄投票实施限制。

Carney faces crucial test after walking away from Trump's deal

The Canadian prime minister will have to sell his gamble that walking away from talks with the White House will be worth the consequences.

中文摘要 加拿大总理卡尼在离开特朗普的交易后面临关键考验,他赌离开白宫谈判将值得后果。

Why is Israel building new illegal settlements?

Israel opens construction bids for new illegal housing units in the occupied West Bank.

中文摘要 以色列在占领的西岸开放新的非法住房单元的建筑投标。

Envoy says Israel did not give the US notice before strikes on Syrian base

US official Tom Barrack has criticised Israel for failing to give his country adequate warning, a claim Tel Aviv denies.

中文摘要 美国官员汤姆·巴拉克批评以色列在袭击叙利亚基地之前没有提前通知美国,特拉维夫否认了这一指控。

Bloomberg Previews Jackson Hole Symposium

Bloomberg's Joe Weisenthal, Lisa Abramowicz, Tracy Alloway and Tom Keene preview the Jackson Hole Symposium as all eyes will be on Fed Chairman Kevin Marsh and his first annual symposium address. (Source: Bloomberg)

中文摘要 彭博新闻前瞻杰克逊霍尔研讨会,美联储主席凯文·马尔什将发表首次年度演讲。

Juicy Yields Draw Junk Bond Buyers to Investment-Grade AI Debt

Companies looking to finance data center projects are increasingly turning to junk bond investors to help them raise billions of dollars, even for debt that is investment-grade.

中文摘要 数据中心项目公司越来越多地寻求垃圾债券投资者帮助筹集数十亿美元,即使债务为投资级。

Bloomberg This Weekend 8/22/2026

The news doesn’t stop when markets close. Hosts David Gura, Christina Ruffini and Alexis Christoforous bring clarity, context and a bit of humor to the weekend’s biggest headlines, LIVE from New York. Joined by Former MI6 Chief Sir Richard Dearlove, American Whiskey Association President & CEO Micha

中文摘要 彭博新闻周末节目聚焦市场收盘后新闻,主持人与英国前MI6负责人等探讨热点。

Pointed! Bloomberg's Weekly News Quiz For Risk-Takers

Pointed offers a strategic twist to the news quiz format, testing not just players’ knowledge of the news but also their confidence in their answers. Join Bloomberg's Alexis Christoforous, David Gura, and Christina Ruffini as they play and check out the quiz for yourself at Bloomberg.com (Source: Bl

中文摘要 彭博新闻每周风险承担者新闻问答,测试玩家对新闻的了解和信心。

US Colleges Face a Shrinking Student Pipeline

US colleges are confronting a prolonged demographic squeeze as the number of high school graduates, after peaking in 2025, is projected to decline 13% through 2041, putting particular pressure on regional schools with limited endowments. Joining David Gura and Christina Ruffini on Bloomberg This Wee

中文摘要 美国大学面临学生数量下降的长期压力,高中生人数预计到2041年将减少13%。

Sanctions Target Iran’s Weakened Regional Network

The US is preparing a new campaign of economic pressure on Iran after diplomacy and months of fighting failed to produce a durable settlement, with President Donald Trump warning countries against maintaining trade ties with Tehran. Atlantic writer and Columbia University Institute of Global Politic

中文摘要 美国对伊朗的新一轮经济制裁,警告国家不要与德黑兰保持贸易关系。

Small Dogs, Big Demand as Dachshund Popularity Takes Off

Dachshunds are enjoying a global popularity boom, climbing to No. 5 in the American Kennel Club’s 2025 rankings as owners embrace their small size and outsized personalities. The surge is supporting everything from breed meetups to dachshund cafes, even as owners contend with health concerns includi

中文摘要 腊肠犬在全球受到热捧,在美国犬业俱乐部排名第五,推动相关产业增长。

Study Shows AI Hitting Paychecks, Not Payrolls

Apollo Chief Economist Torsten Slok says early labor-market data suggest AI’s first impact is appearing more in wages than employment. In a study of hundreds of occupations, Slok found that jobs with higher exposure to AI have seen weaker wage growth, while the employment effect remains relatively s

中文摘要 研究表明,人工智能首先影响的是工资而非就业,高AI接触度工作工资增长较弱。

American Whiskey Pays Price for Canada Trade Rift

US spirits exports to Canada fell 70% from March through December 2025 after most provinces removed American alcohol from their shelves, while exports elsewhere rose 2.5%. American Whiskey Association President & CEO Michael Bilello is on Bloomberg This Weekend and says restoring market access is on

中文摘要 美国威士忌因加拿大贸易分歧付出代价,对加拿大出口下降70%。

MI6 Veteran Warns Against Politicizing Intelligence

Former Head of MI6 Sir Richard Dearlove joins Christina Ruffini on Bloomberg This Weekend to discuss the current state of the international intelligence community. Dearlove says that UK-US intelligence cooperation remains close but expressed concern about politicization and personnel changes inside

中文摘要 前MI6负责人警告不要政治化情报,英国-美国情报合作依然紧密,但存在政治化担忧。

该源今日无内容。