tester-army/e2e
TypeScript · ★ 3,011 · 🍴 114 · 📈 344 stars today
Next generation e2e testing framework for web and mobile apps.
中文介绍 提供新一代的端到端测试框架,适用于Web和移动应用,简化测试流程。
TypeScript · ★ 3,011 · 🍴 114 · 📈 344 stars today
Next generation e2e testing framework for web and mobile apps.
中文介绍 提供新一代的端到端测试框架,适用于Web和移动应用,简化测试流程。
JavaScript · ★ 76,238 · 🍴 4,546 · 📈 1,170 stars today
The design language that makes your AI harness better at design.
中文介绍 设计语言,使AI助手在设计中表现更佳,提升设计效率。
JavaScript · ★ 53,025 · 🍴 7,915 · 📈 270 stars today
Marketing skills for Claude Code and AI agents. CRO, copywriting, SEO, analytics, and growth engineering.
中文介绍 为Claude Code和AI代理提供市场营销技能,包括CRO、文案、SEO、分析和增长工程。
JavaScript · ★ 154,797 · 🍴 8,320 · 📈 1,894 stars today
Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
中文介绍 使AI代理具备最懒散资深开发者的思维模式,减少不必要的代码编写。
Python · ★ 16,835 · 🍴 1,739 · 📈 75 stars today
Give your agent CAD superpowers.
中文介绍 赋予AI代理CAD功能,实现文本到CAD的转换。
Python · ★ 90,807 · 🍴 7,980 · 📈 979 stars today
Give your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
中文介绍 为AI代理提供全网络视野,整合多个平台信息检索功能,无需API费用。
Python · ★ 45,368 · 🍴 4,903 · 📈 152 stars today
Developer-first error tracking and performance monitoring
中文介绍 面向开发者的错误追踪和性能监控工具,帮助开发者快速定位和解决问题。
Python · ★ 63,160 · 🍴 8,055 · 📈 361 stars today
World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant into a full video production studio.
中文介绍 首个开源的视频制作系统,提供多生产流程、工具和技能文件,将AI编程助手转变为视频制作人。
TypeScript · ★ 25,124 · 🍴 6,508 · 📈 492 stars today
中文介绍 提供代码生成工具,具体功能未描述。
Go · ★ 76,527 · 🍴 5,045 · 📈 226 stars today
Fast and extensible multi-platform HTTP/1-2-3 web server with automatic HTTPS
中文介绍 快速且可扩展的多平台HTTP/1-2-3网络服务器,支持自动HTTPS。
JavaScript · ★ 101,190 · 🍴 10,615 · 📈 336 stars today
Production-grade engineering skills for AI coding agents.
中文介绍 为AI编程代理提供生产级工程技能,提升开发效率。
TypeScript · ★ 96,097 · 🍴 8,486 · 📈 627 stars today
Persistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
中文介绍 为每个代理提供跨会话的持久上下文,记录并压缩代理行为,在后续会话中注入相关上下文。
TypeScript · ★ 135,143 · 🍴 20,092 · 📈 121 stars today
Use Garry Tan's exact Claude Code setup: 23 opinionated tools that serve as CEO, Designer, Eng Manager, Release Manager, Doc Engineer, and QA
中文介绍 提供Garry Tan的Claude Code完整设置,包括CEO、设计师、工程师、发布经理、文档工程师和QA等工具。
TypeScript · ★ 92,108 · 🍴 9,082 · 📈 512 stars today
The open-source CapCut alternative
中文介绍 开源的CapCut替代品,提供视频编辑功能。
C · ★ 23,419 · 🍴 2,251 · 📈 211 stars today
DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
中文介绍 DeepSeek 4本地推理引擎,支持Metal、CUDA和ROCm,用于深度学习。
👍 25
Video generation models are increasingly being explored as world simulators for embodied planning and learning. To do so effectively, these models must not only generate visually appealing frames, but also predict how environments dynamically evolve when executing goal-directed actions. While evalua
👍 11
What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahead of them at sensing which prior idea, buried in an ever-growing archive of research, a new problem needs. To study this skill, we draw on researchers who know firsthand which ear
👍 17
On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a frozen or EMA version of themselves that receives privileged information. Its application to multimodal large language models (MLLMs), however, remains larg
👍 39
Looped Transformers achieve parameter efficiency by repeatedly executing a shared block across recurrent loops. Each loop yields an intermediate representation decodable for the same next token, yet standard decoding discards earlier states. Because earlier loops embody less computation, recurrence
👍 11
Generalization in large language models (LLMs) is the ability to produce consistent and semantically stable outputs when the same input is expressed in different ways. Existing work typically evaluates generalization through aggregate accuracy on a single prompt format, task, or set of variations, w
👍 5
Video Large Language Models (VideoLLMs) receive frames in sequential order and interpret how visual content evolves along the temporal axis, yet temporal reasoning remains a persistent weakness across architectures. Reversing the frame order of a video, a transformation that should invert temporal a
👍 45
Vision language model (VLM) agents can control robots through visual feedback and action primitives, but repeated model invocations and redundant observations incur substantial token overhead. We introduce PyRUA-Lean, an interactive code-execution framework that couples feedback-driven primitive com
👍 165
Streaming video LLMs must retain evidence before its relevance to future tasks is known and respond when sufficient evidence becomes available. The challenge is to form reusable factual memory without compromising real-time perception. We introduce OneStreamer, which jointly learns query-independent
👍 11
Recent video generation models can produce highly realistic videos from natural language instructions, with visual quality approaching cinematic standards. Existing evaluation benchmarks, however, predominantly assess visual quality, aesthetic appeal and physical plausibility, while paying limited a
👍 19
Extending a text embedding model to new modalities typically degrades text retrieval quality, and existing omni-modal embedders compensate with multi-billion parameters. We present Omni-Embed-Mini, a 0.9B-parameter model that maps text, speech, audio, images, video, and visually-rich documents into
👍 13
Retrieval-based speculative decoding (SD) drafts tokens by copying continuations from existing text, which suits coding agents that repeatedly reproduce code, logs, and earlier attempts. Yet existing methods fall short in agent pipelines: much of the reusable text is missing from their corpora or st
👍 8
LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable comm
👍 5
Video world models require persistent scene memory to maintain consistency during long-horizon video generation. Existing spatial memories accumulate RGB observations or latent features, increasing storage requirements as generation proceeds. We introduce Honeycomb, a video world model built on HexM
👍 7
LLM agents generate intermediate reasoning and actions token by token, making extended interactions slow and computationally expensive. Jev-style models offer fast probabilistic predictions over finite fields, but require those fields to be specified in advance. This requirement limits autonomous ta
👍 41
Furnished floor plans support real-estate visualization, interior design, and architectural workflows, yet automatic furnishing remains challenged by limited real-world data and the need to satisfy interacting geometric and functional constraints. We ask whether professional furnishing knowledge can
👍 23
Vision-Language Models (VLMs) have shown strong multimodal reasoning capabilities, yet whether they truly capture the physical consistency underlying real-world dynamics remains unclear. Existing benchmark paradigms often suffer from fragmented evaluation, focusing on isolated cognitive stages while
👍 57
Real-world embodied tasks, from everyday activities to professional procedures, require agents to act under physical constraints while tracking evolving object and task states. Tool use sits at the heart of such tasks, as many everyday and professional activities are tool-mediated. Understanding the
👍 16
On-policy self-distillation (OPSD) trains mathematical reasoning models using a privileged teacher that sees a reference solution and supervises student-sampled prefixes. Standard OPSD uses one fixed parameter setting at every state, but nearby settings may offer additional supervision. We find that
👍 7
An assistant that serves the same user over a long horizon has to answer from what that user has revealed: which preferences still hold, which were revised, and which constraints apply now. Retaining that information is not the same as acting on it, and the two are usually optimized as if they were.
👍 12
Coding agents are beginning to move beyond purely digital tasks to tackle physical-world challenges, particularly in robotics. Existing robotics benchmarks, however, primarily focus on the performance of individual artifacts, such as policies or controllers, offering limited coverage of coding agent
👍 13
Backdoor attacks can be implanted in Large Language Models (LLMs) during training, causing unwanted behaviour when a trigger appears in the input. Existing backdoor defences for LLMs attempt to remove the backdoor but inadvertently shift the model's output distribution to benign prompts, which can r
👍 30
Group Relative Policy Optimization (GRPO) is widely used to train reasoning language models, where it computes advantages by centering and normalizing rewards across rollouts of the same prompt. For multiple rewards, GRPO sums the reward components and normalizes the total reward by its within-group
👍 60
Masked diffusion models (MDMs) generate sequences by progressively unmasking several tokens per denoising step, but their reverse process is typically factorized over positions, limiting sample quality in the few-step regime where diffusion's speed advantage over autoregressive decoding matters most
👍 11
Audio-visual large language models (AVLLMs) have made remarkable progress in multimodal understanding and reasoning through interactions among visual, auditory, and linguistic information. However, recent studies show that AVLLMs face a critical challenge: source-confused grounding hallucination, wh
👍 175
On-policy learning has been argued to reduce catastrophic forgetting, produce sparser parameter updates, and improve generalisation. However, existing comparisons between supervised fine-tuning and reinforcement learning vary many factors simultaneously, making the contribution of rollout policy dif
👍 6
Before invoking external tools, an agentic LLM must select among a K-way action space: executing a call, seeking clarification, answering directly, or declining. While internal activation steering can alter these pre-execution decisions, conventional aggregate metrics obscure where altered states la
👍 5
Natural-language service requests can require a language-model decision before execution starts, consuming part of the request's latency budget. We integrate Jev's decision-oriented application programming interface (API) into edge service orchestration to reduce this overhead while retaining servic
👍 56
Multi-step agents are trained on flat action streams: SFT and RLVR weight every token uniformly and ignore the sub-procedures that recur across tasks, the hierarchy that lets humans plan top-down from reusable routines. This structure sits unused, and flat training uses each scarce trajectory less f
👍 13
Multidisciplinary tumor boards integrate multimodal clinical observations and longitudinal patient histories through specialist discussions, yet benchmarks rarely capture these real-world trajectories. We introduce OpenTumorBoard, a benchmark with 611 patient cases and 19,157 discussion turns across
👍 51
Decentralized multi-agent path finding (MAPF) with communication requires agents to reach individual goals without collisions under partial observability. Learnable policies trained on expert data provide an effective approach to this problem. However, when several coordinated joint actions are vali
@UberEng · 66.9K 粉丝 · 54.0K 阅 · 529 赞 · 59 转
Introduction Uber’s rapid adoption of AI agents has fundamentally changed how teams interact with code, data, and operational systems. Early ad-hoc integrations with the MCP (Model Context Protocol)
中文介绍 Uber分享MCP Gateway设计,探讨AI代理如何改变团队与代码、数据和操作系统的互动方式。
@gregisenberg · 724.2K 粉丝 · 44.1K 阅 · 569 赞 · 53 转
For the FIRST TIME in HISTORY, you can describe a physical product in one sentence and, a week later, hold it in your hands, cut from metal and built to your exact measurements. That's VIBE
中文介绍 gregisenberg介绍VIBE制造技术,可实现描述产品后一周内成品到手。
@sermakarevich · 549 粉丝 · 33.0K 阅 · 520 赞 · 64 转
An article for everyone who ships, buys, or signs off on software built on large language models (LLMs): engineers, product managers, and the CEO. Written in plain language, from the big picture down
中文介绍 sermakarevich文章探讨如何判断基于大型语言模型的AI系统是否真正有效。
@VibeMarketer_ · 37.8K 粉丝 · 31.4K 阅 · 504 赞 · 59 转
i'm going to show you how to build a content machine with Opus 5.5 that turns your ideas, research and expertise into articles, posts and newsletters in your own voice. from finding your next topic to
中文介绍 VibeMarketer_教程,展示如何使用Opus 5.5构建内容机器,将想法和专业知识转化为文章。
@onlyonealexia · 1.6K 粉丝 · 26.5K 阅 · 509 赞 · 53 转
A List of Web3, AI & Web2 Hackathons Compiled by @onlyonealexia There are a lot of hackathon lists that look impressive until you actually open the links. I compiled a list of ongoing and upcoming
中文介绍 onlyonealexia汇总Web3、AI和Web2黑客松活动列表,帮助开发者寻找合适的比赛。
@rehan_shei · 11.9K 粉丝 · 24.9K 阅 · 520 赞 · 49 转
A week ago I pointed Opus 5.5 at Terraria before bed. I woke up to a fully working, reverse engineered version of the game. Then I added a homing missile launcher, a tactical nuke and a bunch of new
中文介绍 rehan_shei分享使用Opus 5.5逆向工程Terraria游戏并添加新功能的经历。
@goyal__pramod · 11.4K 粉丝 · 24.7K 阅 · 504 赞 · 59 转
Now I am aware you must have been looking forward to the second part of my CUDA blog (if you haven't read it yet, check it out!), but hey! A man can have varied interests. And I believe if you are
中文介绍 goyal__pramod介绍从零开始学习强化学习的第一个部分。
中文介绍 微软的AI模型在数据库确认前完成了任务。
a quiet day.
中文介绍 今天市场平静无变化。
Learn how startups can choose GPT-6 models, tune reasoning effort, improve prompts and skills, coordinate tools, and prepare workflows for production.
中文介绍 OpenAI发布GPT-6系列模型构建指南,包括模型选择、推理调整、提示改进、工具协调和生产工作流程准备。
Enterprise AI is no longer a future ambition. It is in full operational flight. Model capabilities are advancing faster than most organizations can absorb, while the cost of performance continues to fall. Globally, AI investment is set to reach $2.5 trillion in 2026, up 44% from the previous year. F
中文介绍 企业AI已全面投入运营,预计2026年全球AI投资将达到25万亿美元,同比增长44%。
中文介绍 Hugging Face开源Asta中的快速报告生成模型AstaBrief。
After leading Meta’s Llama models, Ahmad Al-Dahle is now transforming Airbnb with AI — from how its teams develop products to how it serves guests.
中文介绍 前Meta的Llama模型负责人Ahmad Al-Dahle现在用AI改变Airbnb,从产品开发到客户服务。
On an afternoon in Seoul in March 2016, I watched a program I helped build put a stone on the fifth line of a Go board in what looked like a gift to its human opponent. Move 37 in game two of the five-game match looked so absurd that some commentators thought it was a…
中文介绍 不要被误导,大型语言模型并不具备推理能力。
the minimalist harness goes stable... and TypeScript!
中文介绍 Pi 1.0和Pi Durable发布,TypeScript支持上线。
中文介绍 Hugging Face推出AutoSynthData,用于为企业智能体生成训练数据。
We catch up with RLM first author Alex Zhang, MIT PhD, on Jev, PhD masxing, and the future of harnesses.
中文介绍 RLM第一作者、MIT博士Alex Zhang讨论Jev、博士masxing和马具的未来。
Chatham Financial uses Codex and GPT-5.6 to build technology and redesign workflows, cutting trade validation from 30 minutes to under 4.
中文介绍 Chatham Financial利用Codex和GPT-5.6技术,将交易验证时间从30分钟缩短至4分钟以下。
中文介绍 决策模型、Claude-shaped科学、OpenAI安全裁员等AI相关新闻汇总。
Advanced AI may matter most for the routine work behind breakthrough ideas. Explore why execution could shape the next economy and the pace of progress.
中文介绍 高级AI可能对突破性想法背后的日常工作最为重要。
Albertsons Cos. is using ChatGPT Enterprise and the OpenAI API to help teams work faster and make grocery shopping easier for millions of customers.
中文介绍 Albertsons公司利用ChatGPT Enterprise和OpenAI API来提高团队效率和简化顾客的购物体验。
... but you can’t try it yet unless you are “government users and trusted cyber defenders in the Fairwind Program”
中文介绍 Gemini 4 Argon发布,面向政府用户和公平风计划中的受信任网络安全防御者。
Biljana Plavsic, a Bosnian Serb, pleaded guilty to one count of crimes against humanity. She later recanted and said she “would do the same again.”
中文摘要 前波斯尼亚塞尔维亚政治领导人比利扬娜·普拉西奇去世,曾因反人类罪认罪后又撤销
Hundreds of schools will suspend some or all classes on Monday, the education minister said, as students protest conditions in the education system. Thousands of protesters have been arrested and hundreds of people injured.
中文摘要 法国数百所学校将于周一暂停部分或全部课程,因学生抗议教育系统条件,数千人被捕,数百人受伤
Luiz Inácio Lula da Silva is standing against Flávio Bolsonaro in the first round of the presidential race The incumbent, Luiz Inácio Lula da Silva, is up against Flávio Bolsonaro, the son of the former president Jair Bolsonaro, who is in prison for plotting a military coup. The president in Brazil
中文摘要 巴西2026年总统选举首轮投票,卢拉·达席尔瓦对阵前总统之子弗拉维奥·博尔索纳罗
Brazil's environment is a global concern. But on the campaign trail, it's struggling to get voters' attention.
中文摘要 巴西总统选举中,亚马逊的未来虽是全球关注,但在竞选活动中未能吸引选民注意力
Brazil is a global supplier of oil, soybeans and critical minerals. Americas Quarterly editor in chief Brian Winter explains how Brazil's presidential election could impact trade with the U.S. and China.
中文摘要 巴西总统选举可能改变其与美国和中国的关系,巴西是全球石油、大豆和关键矿产的供应商
In Brazil, President Lula da Silva faces Flávio Bolsonaro, son of the jailed former president, in a tight race. If neither wins more than half the votes they'll face each other again on Oct. 25.
中文摘要 巴西总统选举中,卢拉面对博尔索纳罗之子,如无候选人赢得超过半数选票,则将在10月25日进行第二轮投票
New York lawmakers will review sexual-assault laws, including voluntary intoxication rule at the heart of Cornell case.
中文摘要 纽约州将审查性侵犯法律,包括科奈尔案中的自愿醉酒规则
The co-pilot’s embrace of extreme Islamist views prompted Omani officials to scrutinize him, according to two people briefed on the investigation.
中文摘要 飞往迪拜的航班飞行员因极端伊斯兰主义观点受到奥曼官员审查
Republic of Ireland football players again wear black armbands and refuse handshakes with Israel in UEFA Nations League.
中文摘要 爱尔兰在国家队联赛中拒绝与以色列握手并佩戴黑色臂章
If no candidate gets more than 50% of the vote, the election will go to a run-off on 25 October.
中文摘要 巴西投票结束,卢拉和弗拉维奥·博尔索纳罗票数相近,如无候选人得票超过50%,则10月25日将进行第二轮投票
More than 200 U.S. intelligence and military analysts are in Saudi Arabia helping it provide assistance to its Yemeni allies, current and former U.S. officials say.
中文摘要 胡塞武装声称袭击了沙特阿美公司的设施,也门冲突升级,美国在沙特部署200多名分析师协助
More oil is moving through the Strait of Hormuz. What does that mean for Iran’s wartime leverage?
中文摘要 更多石油通过霍尔木兹海峡,对伊朗战时杠杆有何影响
The withdrawal came with unusual speed, following what U.S. officials said were threats of an Iran-based plot against the installation.
中文摘要 美国从英国拉法姆空军基地紧急撤走轰炸机,应对来自伊朗的威胁
Staunchly pro-Russia politician has dominated campaign and inflamed divisions despite not appearing on ballot Bosnian Serb nationalist politician Milorad Dodik has declared victory for his party in an election in Bosnia shaped by economic pressures, the prospect of the Balkan nation joining the Euro
中文摘要 波黑塞尔维亚民族主义者多迪克宣布其党派在波黑选举中获胜,尽管他未出现在选票上
Israel's president declared that the 'antisemitic lie' deliberately endangers Jews and Israelis.
中文摘要 以色列对英国绿党正式将犹太复国主义定义为种族主义表示愤怒
Asian stocks were poised to open higher after softer US jobs data eased pressure on the Federal Reserve to keep raising interest rates. Brent crude rose as Yemen launched a bid to recapture Houthi-controlled areas.
中文摘要 亚洲股市预期走高,因美国就业数据疲软缓解了美联储加息压力,布伦特原油价格上涨。
Acquisition would be French conglomerate’s largest and enhance its products focused on manufacturers
中文摘要 施耐德电气拟以200亿美元收购软件集团PTC,这是法国集团的最大一笔收购,将增强其面向制造商的产品。
Schneider Electric SE is nearing a deal to acquire US-based engineering software developer PTC Inc. for more than $20 billion, according to a source familiar with the matter.
中文摘要 施耐德电气接近以超过200亿美元收购美国工程软件开发商PTC,据知情人士透露。
Interest rate rises, AI anxiety and midterm elections cool animal spirits after record-breaking start to the year
中文摘要 交易放缓威胁到并购热潮提前结束,利率上升、AI焦虑和中期选举在创纪录的开局后冷却了市场情绪。
Dutch group on verge of handing Japanese rival a small victory as it seeks to exit non-core markets ahead of Axalta merger
中文摘要 阿克苏诺贝尔接近将其东南亚装饰漆业务出售给日本涂料,以退出非核心市场,为Axalta合并做准备。
Nobel Peace Prize announced, and another moment of British by-election jeopardy
中文摘要 选举季节以可能是最具争议的结果开始,诺贝尔和平奖揭晓,英国补选陷入困境。
In time it could bring about an equally disruptive glut
中文摘要 世界正面临一场天然气危机,这可能会带来同样破坏性的过剩。
Bloomberg Supreme Court Reporter Greg Stohr is on Bloomberg This Weekend previewing a new term shaped by the court’s growing emergency docket and a major climate case over whether Boulder, Colorado, can pursue state-law claims against oil companies. Speaking with hosts David Gura and Christina Ruffi
中文摘要 最高法院以气候斗争案开启新学期,法院日益增长的紧急案件和涉及科罗拉多州博尔德市是否可以追究石油公司州法律诉讼的气候案件成为焦点。
The news doesn’t stop when markets close. Hosts David Gura, Christina Ruffini and Lisa Mateo bring clarity, context and a bit of humor to the weekend’s biggest headlines, LIVE from New York. Joined by Ambassdor Alexander Yui, Taiwan's Representative to the United States, Kevin Gordon, Schwab, Head o
中文摘要 市场收盘后新闻不停歇,主持人David Gura、Christina Ruffini和Lisa Mateo从纽约带来周末头条新闻的清晰度、背景和一点幽默。
The growing private market for assessments is leading to concerns over misdiagnoses and providers’ soaring revenues
中文摘要 税收资助的ADHD热潮引发了对误诊和供应商收入飙升的担忧。
Retailers reconsidering self-checkout, a global market for elite British sheepdogs, the proliferation of protein-enhanced foods and fit women want to show off their biceps. Join Lisa Mateo, David Gura and Christina Ruffini for a roundup of headlines you may have missed, but gotta see. (Source: Bloom
中文摘要 零售商重新考虑自助结账,英国精英牧羊犬的全球市场,蛋白质增强食品的激增,以及想展示肌肉的健身女性。
Taiwan’s Representative to the US, Alexander Yui, tells Bloomberg This Weekend that Washington has reassured Taipei its Taiwan policy remains unchanged following President Donald Trump’s meeting with Chinese President Xi Jinping, even as a $14 billion US arms package remains stalled. Speaking with h
中文摘要 台湾代表亚历山大·尤伊表示,尽管美国总统特朗普与中国国家主席习近平会面,华盛顿已向台北保证其台湾政策保持不变,同时140亿美元的美国武器包仍保持稳定。
12 回复 · 程序员 节点
42 回复 · 程序员 节点
27 回复 · 程序员 节点
9 回复 · Apple 节点
13 回复 · 程序员 节点
12 回复 · 程序员 节点
17 回复 · 程序员 节点
21 回复 · 程序员 节点
17 回复 · Apple 节点
31 回复 · Apple 节点
该源今日无内容。