每日简报

2026-08-23

← 历史归档

openai/codex

Rust · ★ 113,289 · 🍴 17,357 · 📈 1,978 stars today

Lightweight coding agent that runs in your terminal

中文介绍 轻量级终端运行编码代理,用于代码编写辅助,简化代码执行和解释过程。

mattpocock/skills

Shell · ★ 231,982 · 🍴 19,805 · 📈 2,684 stars today

Skills for Real Engineers. Straight from my .agents directory.

中文介绍 为真实工程师提供技能框架,从个人技能库中提取,适用于实际开发场景。

affaan-m/ECC

JavaScript · ★ 242,165 · 🍴 36,701 · 📈 428 stars today

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

中文介绍 性能优化系统,结合技能、直觉、记忆、安全和研究优先的开发模式,适用于Claude Code等工具。

obra/superpowers

Shell · ★ 276,177 · 🍴 24,698 · 📈 592 stars today

An agentic skills framework & software development methodology that works.

中文介绍 技能框架和软件开发方法论,提供高效的工作流程和代理技能。

Wei-Shaw/sub2api

Go · ★ 38,779 · 🍴 8,039 · 📈 264 stars today

Sub2API 一站式开源中转服务,让 Claude、Openai 、Gemini、Grok订阅统一接入,支持拼车共享,更高效分摊成本,原生工具无缝使用。

中文介绍 一站式开源中转服务,支持拼车共享,实现不同AI服务的统一接入和使用。

makeplane/plane

TypeScript · ★ 57,211 · 🍴 5,442 · 📈 263 stars today

🔥🔥🔥 Open-source Jira, Linear, Monday, and ClickUp alternative. Plane is a modern project management platform to manage tasks, sprints, docs, and triage.

中文介绍 开源的项目管理平台,替代Jira、Linear等,提供任务管理、文档和问题跟踪功能。

n8n-io/n8n

TypeScript · ★ 201,804 · 🍴 60,301 · 📈 202 stars today

Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.

中文介绍 集成了AI功能的公平代码工作流自动化平台,支持可视化构建和自定义代码,适用于多种集成需求。

anthropics/claude-code

Python · ★ 142,521 · 🍴 22,839 · 📈 141 stars today

Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflows - all through natural language commands.

中文介绍 终端内运行的编码工具,理解代码库,通过执行常规任务、解释复杂代码和操作git工作流来加速编码。

AprilNEA/OpenLogi

Rust · ★ 13,902 · 🍴 372 · 📈 959 stars today

⚡️A native, local-first alternative to Logitech Options+, written in Rust 🦀 — remap buttons, DPI, and SmartShift over HID++. No account, no telemetry.

中文介绍 本地优先的Logitech Options+替代品,使用Rust编写,提供按钮重映射、DPI调整和SmartShift功能。

modular/modular

Mojo · ★ 28,838 · 🍴 3,066 · 📈 395 stars today

The Modular Platform (includes MAX & Mojo)

中文介绍 Modular平台,包括MAX和Mojo,提供模块化音乐制作解决方案。

multica-ai/andrej-karpathy-skills

★ 205,288 · 🍴 21,006 · 📈 379 stars today

A single CLAUDE.md file to improve Claude Code behavior, derived from Andrej Karpathy's observations on LLM coding pitfalls.

中文介绍 基于Andrej Karpathy对LLM编码缺陷观察的CLAUDE.md文件,用于改进Claude Code的行为。

mahlernim/google-timeline-visualizer

Kotlin · ★ 2,559 · 🍴 300 · 📈 441 stars today

Visualize your year in travel using your Google Location History (Timeline) data

中文介绍 使用Google位置历史数据可视化年度旅行,提供直观的旅行时间线。

ripienaar/free-for-dev

HTML · ★ 133,887 · 🍴 14,027 · 📈 915 stars today

A list of SaaS, PaaS and IaaS offerings that have free tiers of interest to devops and infradev

中文介绍 列出具有免费层级的SaaS、PaaS和IaaS服务,适用于开发人员和基础设施开发者。

microsoft/TypeScript

Go · ★ 110,533 · 🍴 13,746 · 📈 163 stars today

TypeScript is a superset of JavaScript that compiles to clean JavaScript output.

中文介绍 JavaScript的超集,编译为干净的JavaScript输出,提高开发效率和代码质量。

cursor/plugins

TypeScript · ★ 4,656 · 🍴 381 · 📈 286 stars today

Cursor plugin specification and official plugins

中文介绍 Cursor插件规范和官方插件,提供丰富的插件扩展功能。

PostHog/posthog

Python · ★ 38,607 · 🍴 3,244 · 📈 288 stars today

🦔 PostHog is the leading platform for building self-driving products. Our developer tools – AI observability, analytics, session replay, flags, experiments, error tracking, logs, and more – capture all the context agents need to diagnose problems, uncover opportunities, and ship fixes. Steer it all

中文介绍 PostHog是一个构建自我驱动产品的领先平台,提供AI可观察性、分析、会话回放等开发者工具。

Tencent/AI-Infra-Guard

Python · ★ 5,486 · 🍴 518 · 📈 161 stars today

A full-stack AI Red Teaming platform securing AI ecosystems via Agent Scan, Skills Scan, MCP scan, AI Infra scan and LLM jailbreak evaluation.

中文介绍 全栈AI红队平台,通过代理扫描、技能扫描、MCP扫描、AI基础设施扫描和LLM越狱评估,保障AI生态系统安全。

FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving

👍 14

Long-context modeling is a pivotal capability for Large Language Models, yet the quadratic complexity of attention remains a critical bottleneck, particularly during the compute-intensive prefilling phase. Our previous work, FlashPrefill, mitigates this cost through instantaneous pattern discovery a

SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?

👍 58

Software increasingly functions as part of the scientific instrument itself, making failures in scientific code capable of compromising not only program behavior but also the evidence underlying scientific conclusions. Yet existing evaluations of coding agents largely emphasize aggregate task succes

Repo0: Design-Driven Zero-to-All Code Generation

👍 17

Large language model agents have made substantial progress in code generation, yet most existing systems assume a predefined repository architecture. This assumption does not hold in zero-to-all code generation, where an agent must construct an entire software project directly from natural-language

MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use

👍 31

Memory has become a key component of large language models, enabling them to retain information and learn from long-term interactions. However, existing memory benchmarks mainly evaluate whether information is correctly extracted, stored, and retrieved, while largely overlooking how retrieved memori

EnvHarness: Awakening Static Worlds for Agent Learning

👍 246

LLM agents learn by interacting with environments, yet these environments are hand-built and static: blind to an agent's weaknesses, and quickly left behind as it improves. While recent environment generation methods attempt to address this, they require domain-specific pipelines, rely on expensive

FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis

👍 112

Training terminal agents requires scalable executable supervision, yet synthesizing high-quality terminal tasks remains challenging. Each task couples an instruction, an initialized environment, a reference solution, and an executable verifier; if these artifacts are generated from inconsistent assu

SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents

👍 6

Agent frameworks increasingly package procedural knowledge as skills: instruction files an agent reads on demand, while public libraries now hold thousands of them. Which skill to read has thus become a decision the policy itself makes in the middle of an episode, yet no existing signal trains it. W

Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL

👍 91

Reinforcement learning (RL) has emerged as a powerful approach for improving reasoning in language and vision-language models, yet its strongest successes still depend heavily on ground-truth supervision (e.g., verifiable reward). Such annotations are costly to obtain and become increasingly scarce

Chain-of-Experience for Continual LLM Improvement

👍 7

Humans continuously learn from experience, whereas conventional large language model (LLM) evaluations ignore the models' ability to improve through inference-time interaction. In this paper, we study how LLMs learn from iterative experience at test time, a setting we refer to as Chain-of-Experience

Towards Real-Time and Adaptable LiDAR Scene Completion

👍 3

LiDAR scene completion is a key component of 3D perception in autonomous driving, where the scene must be completed in real time to be usable in downstream tasks. Existing approaches typically follow an initialize-and-refine paradigm, in which a coarse initialization of the scene is first constructe

LLMs Get Smarter from Targeted Synthetic Multilingual Data

👍 4

Language-specific competency (LSC) is the phenomenon of a language model performing better or worse depending on the language of the prompt. In other words, a language model outputs different (and potentially incorrect) responses to the same semantic query when prompted in different languages. Prior

Bounded Agents: Delegation Security for Multi-Agent AI Systems

👍 3

LLM-based agents can act on behalf of a user to access cloud services, call tools, or invoke agents. At session start, the agent's permissions are set but remain static, and each request is evaluated independently, without considering prior actions. Within its permissions, an agent may act contrary

The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning

👍 17

Popular facts are memorised more deeply during pretraining and resist removal longer than rare ones, yet existing LLM unlearning methods apply uniform gradient pressure regardless of training-data frequency. We propose the AdaPop (Adaptive Popularity) method, which combines local token confidence wi

The Embedder's Dilemma: LLMs Are Better, but at What Cost?

👍 7

Should you replace your text-embedding pipeline with a large language model? We answer this with a controlled, cost-aware comparison of ten LLMs across six families and 26 embedding models (118M to 14B parameters) on 37 tasks spanning classification, semantic textual similarity (STS), clustering, pa

QuoteBench: How Matched Scores Can Hide Command-Path Failures

👍 7

LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched execution scores alone cannot distinguish command-generation errors from failures introduced after generation. QuoteBench measures this boundary with exact final-state validation on 5

SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback

👍 29

Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently possess no closed loop through which they might improve from the interaction failures they actually cause. Recent work does close this loop, but derives its feedback from single-turn question-an

OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

👍 88

Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and code execution to manuscript preparation. Yet workflow coverage alone does not provide access to the full evidence on which scientific discovery depend

Temporal Multi-Signal Fusion for Token-Level Hallucination Detection

👍 3

Token-level hallucination detectors score each token independently from a single signal, and fail exactly when the generating model is confidently wrong. This paper instead treats hallucination as a temporally extended span and detects it by sequence labeling: each token is scored from a 33-dimensio

Update on our post-training effort

@gabepereyra · 11.4K 粉丝 · 616.8K 阅 · 507 赞 · 99 转

Authors: @nikogrupen @ItsJulioPereyra @calvincongelado @vtrengarajan @gabepereyra Over the past six months, Harvey’s research agenda has focused on two goals: Building frontier legal intelligence

中文介绍 Harvey团队过去六个月致力于构建前沿法律智能,关注法律情报构建。

How To Set Up Grokbot The Right Way So It Runs Your Business While You Sleep In 2026

@heynavtoor · 149.3K 粉丝 · 161.9K 阅 · 550 赞 · 93 转

Elon Musk just told his followers to try Grokbot after a power user showed it replaced his $10,000 Mac Mini local AI setup (Elon on X, Aug 16). Gavin Baker says his AI usage went up about 100x after

中文介绍 Grokbot效率提升100倍,可替代价值10,000美元的Mac Mini AI系统,适合企业使用。

Building software factories (with no slop)

@dzhng · 13.3K 粉丝 · 137.6K 阅 · 502 赞 · 48 转

The amount of code being written today far outpaces our ability to review it. That part isn't controversial. What's more interesting is that we're visibly losing — the amount of AI slop in modern

中文介绍 软件工厂建设,关注AI代码冗余问题,提高代码审查效率。

Agentic coding is now multiplayer: Introducing Slack Code

@SlackHQ · 434.3K 粉丝 · 86.9K 阅 · 512 赞 · 44 转

AI coding is now a team sport. Code channels bring software development out of private tabs, so teams and agents can write, review, and ship code together, in the open. Slack is where agents become

中文介绍 Slack Code推出,支持多人协作编写、审查和发布代码,促进开源。

AI Engineering Skills Map: Building and Deploying AI Applications

@AndrewYNg · 1.8M 粉丝 · 62.4K 阅 · 1.2K 赞 · 189 转

I previously wrote about our AI Engineering Skills Map, with the highest level skills being (i) Building and deploying AI applications, (ii) Software engineering fundamentals, (iii) Using coding

中文介绍 AI工程技能图谱发布,强调构建和部署AI应用的重要性。

RL Policy Churn

@ID_AA_Carmack · 3.7M 粉丝 · 58.8K 阅 · 577 赞 · 40 转

I have a good theory about this now. The Phenomenon of Policy Churn https://arxiv.org/pdf/2206.00730 The classic exploration method in RL is “epsilon greedy”, where the agent normally takes the “greedy” action

中文介绍 ID_AA_Carmack探讨RL策略退化现象,提出epsilon greedy方法在探索中的作用。

Grok Bot vs. Hermes: They Look Similar. Underneath, They’re Making Very Different Bets.

@themahis · 4.7K 粉丝 · 43.0K 阅 · 515 赞 · 37 转

One is trying to make AI teammates feel like a managed product. The other is turning the agent itself into an open, composable runtime. Research note — August 20, 2026: For this comparison, I

中文介绍 对比Grok Bot与Hermes,分析两种产品在AI队友体验和开放性上的不同。

5 design patterns for long-horizon agent harness

@GoogleCloudTech · 1.3M 粉丝 · 42.9K 阅 · 506 赞 · 87 转

Any agent can look brilliant on a one-shot task. Give it a week of real work and it starts to fall apart. "Long horizon" means the agent keeps going across days and dozens of sessions instead of

中文介绍 提出5种长期目标智能体利用模式,提升智能体在不同场景下的表现。

System Design for Agent Systems (Part 1)

@kmeanskaran · 13.9K 粉丝 · 32.9K 阅 · 518 赞 · 66 转

Put a frontier model into a badly designed agent system and you get a more articulate failure. Nearly all of the engineering sits outside the model, in the scaffolding that has come to be called the

中文介绍 系统设计对智能体系统的重要性,强调工程架构对模型效果的影响。

The Evolution of the Agent Harness

Models keep absorbing the harness into their weights — soon, it will be a harness for human attention rather than for the model.

中文介绍 模型吸收人类注意力的机制逐渐形成,未来可能成为人类注意力的工具。

Simulation: the new Scaling Law — Joon Sung Park, Simile AI

Simile’s CEO about his journey from the viral Generative Agents to creating 8 Billion Digital Twins of every living human... and why it’s gone from fun exploration to very serious business.

中文介绍 Simile AI CEO讲述其从生成代理到创建80亿数字孪生的历程,业务从娱乐探索转向严肃商业。

not much happened today

**Ox Alpha** emerged as a mystery model with strong coding and agentic performance, likely a **Zhipu/GLM-family** model such as **GLM-5.3 Vision**. Analysts suggest its gains come from post-training and infrastructure improvements rather than sheer size, based on the **743B base** of **GLM-5.2** wit

中文介绍 Ox Alpha成为神秘模型,表现优异,分析师认为其增长来自训练后和基础设施改进。

Debates over AI consciousness are a trap

“Runaway” AI, “rogue” agents, and “autonomous” actors—the current rhetoric would have you believe that AI agents are not only awake and aware, but angry at their creators. Prominent tech leaders such as Demis Hassabis, Dario Amodei, and Sam Altman push for regulation of these seemingly “superhuman”

中文介绍 关于AI意识的争论陷入误区,多位科技领袖呼吁对这些“觉醒”的AI实体进行监管。

Unlocking hidden revenue streams with market models

Each day, an airline transports tens of thousands of passengers on hundreds of flights. Often these are not straightforward point-to-point routes, with passengers requiring multiple connections. The airline can consider potentially hundreds of variables to price each of these journeys: demand, seaso

中文介绍 通过市场模型挖掘隐藏的收益潜力。

Introducing AI Futures

Introducing AI Futures, a new OpenAI blog exploring how transformative AI could reshape power, governance, the economy, and individual freedom.

中文介绍 OpenAI推出AI未来博客,探讨AI如何重塑权力、治理、经济和个人自由。

not much happened today

**OpenAI** and **Anthropic** expanded their agent platforms with new desktop features, collaborative editing, and composable APIs like Skills and Files API. **OpenAI** rolled out memory and workflow features in the EEA, UK, and Switzerland. **AT&T** revealed that 40% of employee AI usage routes to o

中文介绍 OpenAI和Anthropic扩展其代理平台,推出新桌面功能、协作编辑和可组合API,AT&T透露40%员工...

Australia news live: Victorian Liberals announce plan for new appeals court to ensure tougher sentences

New court of criminal appeal would give effect to new bail and sentencing laws, opposition leader says. Follow live The Victorian Liberals have announced that, if elected, they will split the court of appeal into two and appoint a new judges to set precedent they say will match “community expectatio

中文摘要 维多利亚州自由党宣布,如当选将设立新的刑事上诉法院,以实施更严格的保释和刑罚法律。

Trump’s Top Trade Representative Details Offer That Canada Rejected

In an interview, Jamieson Greer, President’s Trump’s trade representative, laid out details of what the United States offered to Canada before talks crumbled.

中文摘要 美国贸易代表Greer详细介绍了美国向加拿大提出的贸易协议细节,但谈判破裂。

Open-air cinema brings movie nights back to Khartoum

Families are returning to the movies at an open-air cinema in Khartoum, offering residents a brief escape from war.

中文摘要 开罗露天电影院重新开启,为居民提供逃离战争的短暂时光。

How Two British Historians Made a Smash Hit Podcast

They took on topics as diverse as fall of the Aztecs, the French Revolution, the assassination of Abraham Lincoln, and history’s greatest monkeys. The rest is history.

中文摘要 两位英国历史学家制作了一档关于多元历史主题的爆红播客。

Carney Slams U.S.-Canada Trade Proposal and Vows Retaliation

Prime Minister Mark Carney gave a powerful speech to Canadians on Saturday morning, hours after ordering negotiators to suspend U.S. trade talks despite President Trump’s punishing tariffs.

中文摘要 加拿大总理Carney抨击了美国-加拿大贸易提议,并誓言报复。

Carney faces crucial test after walking away from Trump's deal

The Canadian prime minister will have to sell his gamble that walking away from talks with the White House will be worth the consequences.

中文摘要 加拿大总理Carney在离开特朗普的交易后面临关键考验。

Why is Israel building new illegal settlements?

Israel opens construction bids for new illegal housing units in the occupied West Bank.

中文摘要 以色列在占领的西岸开启建设新非法住房单元的招标。

Bloomberg Previews Jackson Hole Symposium

Bloomberg's Joe Weisenthal, Lisa Abramowicz, Tracy Alloway and Tom Keene preview the Jackson Hole Symposium as all eyes will be on Fed Chairman Kevin Marsh and his first annual symposium address. (Source: Bloomberg)

中文摘要 彭博社前瞻杰克逊霍尔研讨会,美联储主席凯文·马什将发表首次年度演讲。

Juicy Yields Draw Junk Bond Buyers to Investment-Grade AI Debt

Companies looking to finance data center projects are increasingly turning to junk bond investors to help them raise billions of dollars, even for debt that is investment-grade.

中文摘要 数据中心项目融资公司越来越多地转向垃圾债券投资者,即使债务为投资级。

Bloomberg This Weekend 8/22/2026

The news doesn’t stop when markets close. Hosts David Gura, Christina Ruffini and Alexis Christoforous bring clarity, context and a bit of humor to the weekend’s biggest headlines, LIVE from New York. Joined by Former MI6 Chief Sir Richard Dearlove, American Whiskey Association President & CEO Micha

中文摘要 彭博社周末新闻节目,聚焦市场收盘后的重大新闻。

Pointed! Bloomberg's Weekly News Quiz For Risk-Takers

Pointed offers a strategic twist to the news quiz format, testing not just players’ knowledge of the news but also their confidence in their answers. Join Bloomberg's Alexis Christoforous, David Gura, and Christina Ruffini as they play and check out the quiz for yourself at Bloomberg.com (Source: Bl

中文摘要 彭博社推出针对风险承担者的新闻问答节目,测试玩家的知识和信心。

US Colleges Face a Shrinking Student Pipeline

US colleges are confronting a prolonged demographic squeeze as the number of high school graduates, after peaking in 2025, is projected to decline 13% through 2041, putting particular pressure on regional schools with limited endowments. Joining David Gura and Christina Ruffini on Bloomberg This Wee

中文摘要 美国大学面临学生人数减少的长期压力,预计到2041年将下降13%。

Sanctions Target Iran’s Weakened Regional Network

The US is preparing a new campaign of economic pressure on Iran after diplomacy and months of fighting failed to produce a durable settlement, with President Donald Trump warning countries against maintaining trade ties with Tehran. Atlantic writer and Columbia University Institute of Global Politic

中文摘要 美国准备对伊朗实施新的经济压力,警告各国不要与德黑兰保持贸易关系。

Small Dogs, Big Demand as Dachshund Popularity Takes Off

Dachshunds are enjoying a global popularity boom, climbing to No. 5 in the American Kennel Club’s 2025 rankings as owners embrace their small size and outsized personalities. The surge is supporting everything from breed meetups to dachshund cafes, even as owners contend with health concerns includi

中文摘要 腊肠犬在全球范围内受到欢迎,排名升至美国犬业俱乐部2025年第五位。

Study Shows AI Hitting Paychecks, Not Payrolls

Apollo Chief Economist Torsten Slok says early labor-market data suggest AI’s first impact is appearing more in wages than employment. In a study of hundreds of occupations, Slok found that jobs with higher exposure to AI have seen weaker wage growth, while the employment effect remains relatively s

中文摘要 AI对工资的影响大于就业,高AI接触度的职业工资增长较弱。

American Whiskey Pays Price for Canada Trade Rift

US spirits exports to Canada fell 70% from March through December 2025 after most provinces removed American alcohol from their shelves, while exports elsewhere rose 2.5%. American Whiskey Association President & CEO Michael Bilello is on Bloomberg This Weekend and says restoring market access is on

中文摘要 美国威士忌因加拿大贸易分歧而受损,对加拿大出口下降70%。

MI6 Veteran Warns Against Politicizing Intelligence

Former Head of MI6 Sir Richard Dearlove joins Christina Ruffini on Bloomberg This Weekend to discuss the current state of the international intelligence community. Dearlove says that UK-US intelligence cooperation remains close but expressed concern about politicization and personnel changes inside

中文摘要 前MI6负责人警告不要政治化情报,英国-美国情报合作保持紧密,但担忧政治化。

该源今日无内容。