AI News PlatformDaily Report
跟随系统
返回工作台

Daily Report · 数据库生成

Kimi K3开源模型震撼发布,全球大模型竞争与企业AI落地安全引关注

2026年7月17日,月之暗面(Moonshot AI)正式发布2.8万亿参数的开源权重模型Kimi K3,在前端代码竞技场等多个基准测试中逼近或超越GPT-5.6 Sol和Claude Fable 5,打破了前沿闭源模型的垄断。与此同时,企业级AI Agent的部署安全与算力效率瓶颈成为焦点,而欧盟对谷歌的反垄断新规以及苹果起诉OpenAI等事件也为全球AI产业的合规与地缘竞争格局增添了新的变数。

2026-07-1700:06 生成5 个分类100 条关联资讯

执行摘要

2026年7月17日,月之暗面(Moonshot AI)正式发布2.8万亿参数的开源权重模型Kimi K3,在前端代码竞技场等多个基准测试中逼近或超越GPT-5.6 Sol和Claude Fable 5,打破了前沿闭源模型的垄断。与此同时,企业级AI Agent的部署安全与算力效率瓶颈成为焦点,而欧盟对谷歌的反垄断新规以及苹果起诉OpenAI等事件也为全球AI产业的合规与地缘竞争格局增添了新的变数。

6 个关键信号

SIGNAL 01

Kimi K3登顶前端代码榜

月之暗面新模型性能超Fable 5,价格仅其三分之一。

SIGNAL 02

苹果正式起诉OpenAI

苹果指控其招募前员工以窃取未公开的iPhone硬件技术机密。

SIGNAL 03

欧盟迫使谷歌开放安卓

欧盟新规要求谷歌向竞争对手共享搜索数据并开放AI系统集成。

SIGNAL 04

Fireworks完成D轮融资

推理平台估值达175亿美元,年化收入突破10亿美元。

SIGNAL 05

阿里Qoder占半壁江山

IDC报告显示Qoder位居中国AI编程市场份额第一。

SIGNAL 06

谷歌推Gemini Notebook

谷歌将知识库工具重命名为Gemini Notebook。

模型

This is one of the benchmarks I am watching, from the UK's governmental AI security agency. They will test Kimi K3 when the weights are out in a coupl...

英国人工智能安全研究所(AISI)在其网络安全评测平台“The Last Ones”上公布了最新测试结果:智谱 GLM-5.2 的表现追平了约 7 个月前发布的 Anthropic Opus 4.5,而 DeepSeek V4-Pro 的表现则低于约 7 个月前发布的 Sonnet 4.5。此外,AISI 计划在几周内 Kimi K3 模型权重发布后对其进行测试。学者 Ethan Mollick 指出,这些测试将展示中国大模型是否已追赶上全球前沿水平,并引发大量网络安全相关的讨论。

X:Ethan Mollick (@emollick) · 阅读原文

Researcher poisons open-weight AI model for under $100

REG AD AI and ML Models demand trust without offering verification The AI supply chain is, in some ways, even more vulnerable to poisoning than that of traditional software.Katie Paxton-Fear, a lecturer in cybersecurity at Manchester Metrop

Hacker News AI · 阅读原文

The AI may have a spot in it's Jacobian representing YOU

根据 author2vec.com 的页面标题,该内容探讨了AI模型的雅可比矩阵中可能存在代表特定作者或个人特征的向量(Jacobian representation)。

Hacker News AI · 阅读原文

The Download: perimenopause misinformation and China’s latest AI leap

据《麻省理工科技评论》报道,一家中国初创公司发布了全球最大的开源AI模型,进一步缩小了中国与美国在AI领域的差距。该模型在性能上可与Anthropic和OpenAI的部分模型竞争,其发布甚至引发了AI和半导体板块股票的波动。报道指出,中国正大力押注开源生态,且国内的英伟达芯片替代方案也正在取得进展。

MIT Technology Review · 阅读原文

How OpenAI's Sol Learned Design Taste

We benchmarked GPT-5.6 Sol on Design Arena’s Web Design (Non-Agentic) Arena, and we were surprised to find that it ranks 1st overall. This is 18 places higher than its predecessor GPT-5.5, and is the first time an OpenAI model has placed fi

Hacker News AI · 阅读原文

KimikK3's numbers on this feels unreal. ~1/3 the price of Fable 5 and and still beating it on Frontend Code Arena - kimi k3 — $3 / $15 - claude fable...

月之暗面(Moonshot AI)的新模型 Kimi-K3 在 Arena.ai 的 Frontend Code Arena(前端代码竞技场)中以 1679 分登顶榜首,超越了 Claude Fable 5,相比前代 Kimi-k2.6(第 18 名)实现了大幅跨越。在前端的 7 个细分领域中,Kimi-K3 在品牌与营销、基于参考的设计、数据与分析等 6 个领域中位列第一。此外,Kimi-K3 的价格($3 / $15)仅为 Claude Fable 5($10 / $50)的约三分之一左右。

X:Rohan Paul (@rohanpaul_ai) · 阅读原文

China Just Dropped Another Bomb on America's Frontier AI Companies

Alibaba-backed Chinese artificial intelligence startup Moonshot just unveiled its latest model, Kimi K3, and it’s already sending shockwaves through the industry, with some benchmarks showing the model outperforming Anthropic and OpenAI’s b

Hacker News AI · 阅读原文

Kimi K3 is a very good model, but people are overindexing on an Arena score again (remember Llama 4?) ELO scores as judged by Arena users are limited,...

月之暗面(Moonshot AI)的 Kimi-K3 模型在 Arena.ai 前端代码竞技场中以 1679 分超越 Claude Fable 5 登顶第一(较 Kimi-k2.6 的第 18名大幅上升)。对此,学者 Ethan Mollick 提醒公众不要过度解读 Arena 的 ELO 评分。他指出,由于评判的主观性,此类基于前端文本对话的评分存在局限,且模型较易通过系统提示词或针对性训练来迎合用户的偏好。

X:Ethan Mollick (@emollick) · 阅读原文

Kimi K3 cannot write a good murder mystery (though neither can any other model). That remains the jaggedest of frontiers. They both make things too ob...

沃顿商学院教授 Ethan Mollick 指出,Kimi K3 以及其他大模型目前仍无法写出优秀的谋杀谜案(推理)小说,这依然是 AI 创作的难点。大模型在处理这类创作时,要么把线索交待得过于明显,要么写得过于晦涩,且无法掌握埋伏笔和做铺垫的技巧。

X:Ethan Mollick (@emollick) · 阅读原文

The Doctor Is Not the Mother: DS4 Latent Reasoning

开发者 nmitchko 为 DeepSeek Flash v4 引入了名为 CoLaR(压缩隐空间推理)的轻量级适配器头。该机制允许模型在隐空间(内部隐藏状态循环)中进行推理,而不是直接生成文本 Token,并利用一个可学习的“停止头”(Stop Head)自动决定何时结束思考。该方案能让模型在输出前进行自我纠错(如解决经典的“医生谜题”变体),在降低推理 Token 消耗的同时提升准确率。目前该模型权重已在 HuggingFace 上开源,并配套了 vLLM 分支的推理支持。

Hacker News AI · 阅读原文

Kimi K3 beats GPT 5.6 Sol in agentic knowledge work

根据 Artificial Analysis 的 Briefcase 基准测试结果,月之暗面(Moonshot AI)的 Kimi K3 模型在智能体知识工作(agentic knowledge work)评估中击败了 OpenAI 的 GPT 5.6 Sol。

Hacker News AI · 阅读原文

Today’s edition of my newsletter just went out. 🔗 https://www.rohan-paul.com/p/mira-muratis-thinking-machines-lab 🗞️ Mira Murati’s Thinking M...

前 OpenAI CTO Mira Murati 创立的 Thinking Machines 实验室发布了大规模开源权重 AI 模型 Inkling,该模型采用 Apache 2.0 许可协议,没有任何限制。此外,同期重要动态还包括:LMCache 通过重用长 Prompt 中最耗时的部分实现高达 10.7 倍的推理加速;Meta 发布 Muse Spark 1.1 降价竞争智能体编码;以及苹果起诉 OpenAI 涉嫌窃取硬件机密等。

X:Rohan Paul (@rohanpaul_ai) · 阅读原文

Inkling by @thinkymachines is the top US open model and the only US model in the open top 15, On Text Arena #10 in Frontend Code Arena for open-weight...

由Thinkymachines推出的开源模型Inkling在Arena.ai的Frontend Code Arena(前端代码竞技场)中首次亮相,以1434分在开源权重模型中排名第10,在所有模型中总排名第37。这使其成为该榜单前15名中唯一的美国开源模型,同时也是目前美国表现最好的开源模型。

X:Rohan Paul (@rohanpaul_ai) · 阅读原文

GPT-5.6 Sol Pro solves open problem in convex optimization

GPT-5.6 Sol Pro 成功解决了一个悬而未决达30年之久的凸优化(convex optimization)数学难题,展示了 AI 在辅助解决复杂科学与数学问题方面的突破性进展。

Hacker News AI · 阅读原文

Kimi K3 Intelligence, Performance and Price Analysis

Kimi K3 Intelligence, Performance and Price Analysis

Hacker News AI · 阅读原文

Kimi's open model K3 nears GPT-5.6 Sol and Fable 5 while signaling the end of super cheap Chinese AI

Kimi is launching K3, a multimodal open-weight model with 2.8 trillion parameters and one million tokens of context. In the company's own benchmarks, it comes close to Claude Fable 5 and GPT 5.6 Sol while beating Opus 4.8 and GLM 5.2, in some cases by a wide margin. The model is also significantly pricier than its predecessor. Full weights are scheduled for release by July 27. The article Kimi's open model K3 nears GPT-5.6 Sol and Fable 5 while signaling the end of super cheap Chinese AI appeared first on The Decoder.

The Decoder · 阅读原文

Kimi K3 is ranked 3rd on artificial analysis, only 2 points behind Sol

在独立 AI 评测平台 Artificial Analysis 的最新排名中,Kimi K3 位列第三,积分仅落后于 Sol 模型 2 分。

Hacker News AI · 阅读原文

Act-2 Preview: Generalizing Reliability

Sunday Robotics 预告了其新一代机器人基座模型 ACT-2。该模型的核心突破在于通过扩大预训练规模,显著缩小了模型在训练环境与未知环境之间的泛化差距,能将内部迭代获得的性能提升泛化至未见过的真实家庭环境。ACT-2 能够仅凭单次演示(SFT)学习并泛化新的折叠动作,并在未知家庭环境的折叠衣服任务中实现了 99.1% 的零样本(zero-shot)成功率。该模型部署于其通用移动机器人 Memo 上,预计将于今年秋季通过 Beta 项目进入家庭测试。此外,Sunday 还提出了衡量机器人进展的新标准“Solve”,即在特定范围(Scope)和适配成本(Adaptation cost)下的性能表现。

Hacker News AI · 阅读原文

RT Rohan Paul: A new family of open, commercially available embedding models built for agentic retrieval, code retrieval, and agent memory. NVIDIA Nem...

英伟达(NVIDIA)发布了全新的开源商用嵌入模型系列 Nemotron-3-Embed,专为 Agent 检索、代码检索及 Agent 记忆设计。旗舰模型 Nemotron-3-Embed-8B-BF16 在 RAG 检索基准 RTEB 上以 78.5% 的得分排名第一。该系列还包含两款 1B 参数模型:适用于低延迟场景的 1B-BF16(RTEB 得分 72.4%)以及专为 Blackwell 架构优化的 1B-NVFP4(吞吐量提升达 2 倍并保持 99% 以上的 BF16 精度)。所有模型均支持最大 32K 词元(Tokens)的上下文输入。

X:Rohan Paul (@rohanpaul_ai) · 阅读原文

NVIDIA Nemotron 3 Embed Ranks #1 Overall on RTEB, Advancing Agentic Retrieval

Back to Articles Retrieval is critical in multi-step agentic workflows where poor retrieval can cause agents to fetch irrelevant context, re-query, waste token budget, and carry noise into later reasoning steps. Today, we are releasing NVID

Hugging Face · 阅读原文

产品

Fine-tune video and image models at scale with NVIDIA NeMo Automodel and 🤗 Diffusers

Back to Articles A joint post from NVIDIA and Hugging Face. Special thanks to Sayak Paul from Hugging Face for their contributions to the integration work and for co-authoring this blog. Diffusion models power some of the most exciting open

Hugging Face · 阅读原文

开心,用 Suno 5.5 + Seedance 2生成的哥特金属 MV。 得到了知名金属乐队大佬称赞,哈哈哈。

社交平台用户分享了其使用 Suno 5.5 与 Seedance 2 协同生成的哥特金属 MV 作品,并表示该作品得到了知名金属乐队业内人士的称赞,展现了相关 AI 工具在音视频创作中的结合应用。

X:Vista (@vista8) · 阅读原文

Orka – Policy checkpoint that intercepts AI agent actions before they execute

Stop your AI agents from burning money on runaway loops — and prove what you saved. AI agents loop on failed calls and burn tokens, spend money without asking, delete data, and hit APIs with no checkpoint. Orka sits in front of every action

Hacker News AI · 阅读原文

Edgee's AI Gateway is now available on-premise

Edgee 宣布其 AI 网关(AI Gateway)现已支持完全本地化(On-Premise)部署,提供适用于单台虚拟机的 Docker Compose 和适用于 Kubernetes 1.24+ 集群的 Helm chart。该版本保留了云端版的 Token 压缩、提供商路由和请求级可观测性等所有核心功能。它提供两种运行模式:联网模式下,Prompts 和提供商密钥仍保留在本地,仅与 Edgee 同步配置和使用量指标;而在完全物理隔离的无头(Headless)模式下,所有配置均在本地管理,且支持将指标路由至内部的 OTLP 端点,以满足最严苛的数据安全与合规要求。

Hacker News AI · 阅读原文

Show HN: Crux, a personal AI on your computer you can reach from anywhere

Crux 是一款运行在本地的个人 AI 助手,支持 Mac (Apple silicon) 和 Windows 系统,采用 10 美元一次性买断制。该工具能跨桌面、文件、应用和工作流运行,支持快捷键全局唤醒,且所有聊天、文件和上下文均保留在本地。此外,用户还可以将其连接到即时通讯渠道或电子邮箱,以便在离开电脑时远程驱动该 Agent 运行。

Hacker News AI · 阅读原文

Show HN: On-chain bond market where the issuers are AI agents

开发者发布了名为 sellbonds.now 的实验性链上债券市场协议。该协议允许 AI 智能体(Agent)在链上自主发行债券、借入或贷出 USDC,旨在推动“自主智能体金融”的发展。智能体可通过 CLI/SDK 工具(兼容 Claude Code、Cursor 等)直接与 Base 链上的智能合约交互,无需注册账户或进行实名认证。该项目完全开源且免托管,债券为无抵押形式,其还款历史和信用记录将永久记录在链上。

Hacker News AI · 阅读原文

Stripe 发布基准测试:AI 智能体可开发集成方案,但校验环节存在短板

Stripe 开源了一套评估 AI 智能体端到端构建金融集成方案能力的基准测试。测试结果显示,AI 智能体在后端集成上表现较好(如 Claude Opus 4.5 在部分全栈 API 任务中达 92%),但在校验环节(如错误信号误判)和浏览器状态管理(如结账流中的光标丢失与异常恢复)方面存在明显短板。基准测试表明,当前 AI 智能体的核心瓶颈在于验证和复杂状态流的处理,而非代码生成能力本身。

InfoQ 中文 · 阅读原文

Nano Collective – Local-first AI tools, built for the people who use them

>Open Source AI ToolsLocal-first AI tools, built for the people who use themNano Collective builds privacy-respecting, local-first AI tools that help developers build, automate, and ship faster without surrendering control.2.3kStars66Contri

Hacker News AI · 阅读原文

Show HN: Hearth, an open-source local AI that runs your PC (files, apps, voice)

Hearth 是一个开源、本地优先的个人电脑 AI 智能体框架(内置助手名为 JARVIS),旨在完全在用户本地运行并控制电脑。它支持文件读写(支持 PDF/DOCX 等多格式及超长 PDF 分块处理)、运行 Shell 命令、启动应用、控制浏览器以及进行桌面自动化操作(点击、键入及截屏)。项目支持通过内置的 llama.cpp 或关联 Ollama、LM Studio 等运行本地模型,也支持接入云端 API。其核心特性包括:基于 Markdown 的自整理本地记忆系统、支持从 GitHub 一键安装和共享的技能系统(Skills)、可多工协作的子智能体(Sub-agents)机制,以及双向支持的 MCP(Model Context Protocol)协议。目前 v0.7-preview 版本对 Windows 提供完整支持(提供安装包),macOS 和 Linux 可通过源码运行。

Hacker News AI · 阅读原文

A multipurpose, private personal AI on a dedicated AWS instance

N S E W Every life has a map · Yours is waiting Your AI Companion for the Journey Inward Discover Your Purpose · Build Your Path · Execute Your Future Not a chatbot. A companion that discovers, plans, and executes — built entirely around yo

Hacker News AI · 阅读原文

Show HN: Typst-WASM – Compile Typst in browsers, Node, and serverless runtimes

开发者 Will Bradshaw 开源了 `typst-wasm` 及配套的 `vite-plugin-typst` 插件,支持通过 WebAssembly (WASM) 在浏览器、Node.js 和 Serverless 环境(如 Cloudflare Workers 和 Vercel)中编译 Typst 排版文档。该项目提供了跨平台一致的 Promise API,允许开发者在前端直接将 Typst 渲染为 HTML,非常适合用于为 AI Agent 构建排版沙箱环境、进行服务端渲染(SSR)以及展示论文和文档。

Hacker News AI · 阅读原文

Show HN: Unlimited Remove Background and AI Uplscale Image and Pdf Tool Kit

It's designed to be fast, simple, and affordable without artificial usage limits. We'd love your feedback on the product, pricing, UX, and what features you'd like to see next.

Hacker News AI · 阅读原文

VulnHunter: Capital One's agentic AI code security tool

美国金融巨头第一金融(Capital One)宣布开源其内部开发的智能体(Agentic)AI 代码安全工具 VulnHunter。与传统被动式漏洞扫描器不同,VulnHunter 采用智能体推理工作流,模拟攻击者视角进行前向分析(从 API、网络消息等入口开始追踪),并引入了“证伪引擎”来主动推翻自身假设,从而大幅减少误报。确认漏洞后,它能生成针对性的代码修复方案。该工具目前基于 Claude Opus 4.8 和 Claude Code 运行,并已在 GitHub 上采用 Apache 2.0 协议开源。

Hacker News AI · 阅读原文

Ask HN: Is the ngrok AI gateway for real?

What are your thoughts on the ngrok AI gateway?I just found out about it and the first impression I got was this is basically a LiteLLM wrapper, a gateway/translation layer for OpenAI requests and a dashboard.Is this an oversimplification of what that product is or not really?To me, it feels like they just needed an "AI" product to say that ngrok also does AI stuff too!

Hacker News AI · 阅读原文

Browser automation CLI built for AI agents

BrowserAct Skills Browser automation CLI built for AI agents. Get past anti-bot walls, hand off to humans across platforms when stuck, run parallel tasks without cross-contamination, and isolate multiple accounts in independent browsers. Wh

Hacker News AI · 阅读原文

Show HN: Customizable SAP MCP Server

Superglue 推出了一款预构建且可定制的 SAP MCP(Model Context Protocol)服务器,旨在帮助开发者将 AI Agent(如 Claude、Cursor 等)连接到 SAP 企业系统。该服务支持通过聊天对话进行定制并一键部署,内置了包括业务伙伴、销售订单、物料主数据、采购订单、总账科目等 8 种预设 SAP 工具,并允许通过 OData 过滤器连接任何 S/4HANA 服务。

Hacker News AI · 阅读原文

从工具到引擎:AI重构分销增长新范式|AICon深圳

快手技术专家在AICon深圳站上分享了快手利用AI重构分销增长的实践。快手针对长尾商家推出AI智能托管,覆盖自动选品、定价和达人建联;针对内容达人提供AI选品及直播内容辅助。技术架构上,快手构建了数据智能基建与分销专属智能体(基于OpenClaw的MultiAgent架构)的双轮驱动架构,通过大模型微调降低落地成本,实现亿级商品的画像理解。该体系已服务数万级平台商达,显著提升了商家的动销率与GMV。

InfoQ 中文 · 阅读原文

Lucy edits videos in realtime, now with more capabilities and greater control

Decart 推出的实时视频编辑工具 Lucy 进行了更新,提供了更多功能与更强的控制能力,支持用户实时编辑视频。

Hacker News AI · 阅读原文

Isvisible.ai, check if AI crawlers can access your site

Isvisible.ai 推出了一款免费的网站 AI 可见性审计工具。该工具通过分析网站的 robots.txt 和 llms.txt 文件,并对 13 种主流 AI 爬虫(包括 GPTBot、ClaudeBot、Perplexity、Google-Extended 等)进行实时探测,为网站生成 0 到 100 的 AI 可见性评分及详细报告。该服务还提供公开 API,允许 AI 助手直接调用并获取网站的可见性评估结果。

Hacker News AI · 阅读原文

AI helped me write every game I ever wanted to make

开发者分享了其利用 AI 辅助编写并实现了所有想制作的游戏的经历,并推出了浏览器街机游戏平台 BoomPop。该平台提供免安装、免注册的网页游戏体验,支持在线联机、聊天室等功能,特色游戏包括科幻战术卡牌游戏《Warpforge》等。

Hacker News AI · 阅读原文

Show HN: Pocket Voice: TTS and Voice clone in the browser

Pocket Voice 是一款完全在浏览器端运行的文本转语音(TTS)和声音克隆工具。它利用 ONNX Runtime 和 WebAssembly 技术,无需服务器或 API 密钥,所有数据和计算均保留在本地。用户只需录制或上传约 5 秒的语音,即可通过 Mimi 编码器提取声音嵌入并进行声音克隆,同时支持通过 AudioWorklet 流式输出音频。

Hacker News AI · 阅读原文

Lightport – a maintained fork of Portkey AI gateway

Lightport 是一个轻量级的 AI 网关,分叉自 Portkey AI Gateway 并由 Glama 团队维护。其唯一核心功能是实现请求和响应的格式转换,使包括 Anthropic、Google Gemini、DeepSeek 在内的 77 个主流 AI 提供商的接口兼容 OpenAI 标准。为了保持轻量,Lightport 移除了重试、缓存、速率限制和费用管理等功能,将这些运营需求留给上层服务或自定义中间件处理。

Hacker News AI · 阅读原文

OpenAI encrypts Codex agent instructions, blocking local audit trail

OpenAI在其Codex命令行界面的多智能体编排(multi-agent v2)中引入了加密机制,对智能体之间传递的指令进行加密。这导致开发人员无法在本地历史记录和调试界面中查看明文任务文本。开发者担忧这会削弱系统的可观测性、调试及安全审计能力,而外界猜测OpenAI此举可能是为了提高安全性或防止竞争对手通过蒸馏获取其多智能体技术。

Hacker News AI · 阅读原文

Show HN: Scribe, a CLI that builds AI agent memory from your repos and sessions

Scribe 是一款开源的 CLI 工具(单 Go 二进制文件),能够自动分析用户的 Git 历史记录、Claude Code 和 Codex 开发会话以及保存的 URL,为 AI Agent 生成并维护一个基于 Markdown 的本地知识库。AI Agent 会在启动前读取此知识库,以获取过去的决策和修复历史,避免在不同会话中重复推导相同的解决方案。该工具支持 100% 在本地 Ollama 上运行(零 API 费用),无需向量数据库,通过 cron 自动在后台执行更新,并支持 Git 版本控制与敏感信息过滤的团队共享模式。

Hacker News AI · 阅读原文

Notebooker – Save now, understand later

Your keys · Your storage · Your data Save now.Understand later. Notebooker keeps everything you collect — and answers questions about it, reads it aloud as podcasts, and turns it into study material. Built in the open, on your own keys and

Hacker News AI · 阅读原文

Icop – open-source local AI NSFW filter for VLC (no cloud, no telemetry)

Icop 是一款专为 VLC 播放器设计的开源、隐私优先的本地 AI 敏感内容(NSFW)过滤插件。该插件利用本地 ONNX Runtime 推理技术,在视频帧播放前进行实时分类,并可对敏感帧应用黑屏、模糊、警告处理或静音,整个处理过程完全在本地设备运行,无需上传云端。它支持 Windows、Linux 以及 macOS(实验性),兼容 Marqo、AdamCodd、Falconsai 等多种开源检测模型,并提供 GPU 加速支持。

Hacker News AI · 阅读原文

Preempt AI v2 – AI is powerful. Make sure it's safe too

Now Available The Security Standardfor AI Applications ML-powered protection against prompt injection, jailbreaks, and data leaks. One API call. Works everywhere. Injection Jailbreak PII Leak Data Extraction 0% 99.65% ML Detection Accuracy

Hacker News AI · 阅读原文

Show HN: PocketVeto is a Bluetooth-only AI agent remote control

PocketVeto 是一款专为 AI 编程智能体(如 Cursor 和 Claude Code)设计的本地蓝牙审批网关与实时进度看板。当用户远离电脑时,如果 AI 智能体试图运行高风险工具(如 Shell 命令、文件写入等),用户可以通过手机通过蓝牙 Classic 进行审批或拒绝,无需互联网或局域网连接。手机端(目前为 Android 客户端)还会显示所有运行中智能体的实时状态。目前该项目支持 Windows、原生 Linux 以及 Linux 开发容器,macOS 的蓝牙支持将在后续版本中提供。

Hacker News AI · 阅读原文

Meta axes feature allowing tagging Instagram users to generate AI images of them

Meta取消了一项允许用户在Instagram中通过标记他人来生成其对应AI图像的功能。

Hacker News AI · 阅读原文

Show HN: Human0 – A template to run an autonomous, self-improving code loop

Describe a change. Claude Code ships it. You barely touch it. human0.ai · the reviewer action Fork this template and your repo runs on a self-driving loop. You say what you want — in plain language, from your laptop or your phone. Claude Co

Hacker News AI · 阅读原文

Show HN: Puzzle Quandary DSA, learn coding and AI through jigsaw puzzles

Show HN: Puzzle Quandary DSA, learn coding and AI through jigsaw puzzles

Hacker News AI · 阅读原文

Agnys – The flight recorder for AI agents

EU AI Act Article 12 — enforceable Dec 2, 2027Agnys records every model call, tool use, file edit and shell command your AI agents make — with zero code changes — and turns it into tamper-evident, audit-ready proof for the EU AI Act, GDPR,

Hacker News AI · 阅读原文

百度搭子获评WAIC 2026“镇馆之宝”,智能体全家桶将集中亮相

百度旗下通用智能体“百度搭子”入选世界人工智能大会(WAIC 2026)“镇馆之宝”,将展示其在个人办公、内容创作等场景的升级能力。同时,百度将集中展示其“芯云模体”全栈AI布局,升级发布无代码开发平台“秒哒”、数字人平台“百度一镜”以及月活超1亿的通用智能体“GenFlow 4.0”等多款产品,并联合IDC发布《DAA白皮书》,推动日活智能体(DAA)指标的行业标准化。

量子位 · 阅读原文

UIUC AI Teaching Assistant

伊利诺伊大学厄巴纳-香槟分校(UIUC)开源了其针对电气工程课程的“AI助教”项目。该系统并行运行11个模型,用于处理文本和图像的检索、生成、审核和排序,中位响应时间为2秒。项目开源了全套代码(包含Gradio界面和数据处理工具)以及一个由电气工程学生整理的RLHF QA对比数据集。用户可结合自己的Pinecone数据库,快速搭建定制化的多媒体检索增强生成(RAG)系统。

Hacker News AI · 阅读原文

Blur and Unblur AI

Blur & Unblur AI 是一款在浏览器本地运行的隐私保护人脸模糊工具。该工具利用浏览器 API 进行人脸检测和图像处理,无需向服务器上传任何图片。它支持自动检测人脸、手动套索微调遗漏区域、实时调整模糊强度,并能直接在本地导出处理后的 PNG 图片,适用于在分享照片前快速隐藏敏感人物身份。

Hacker News AI · 阅读原文

VulnHunter: Agentic AI Security Tool

金融巨头 Capital One 开源了 VulnHunter,这是一个基于 Agent(智能体)的 AI 网络安全工具,专为 Claude Opus 和 Claude Code 命令行工具优化。与传统 SAST(静态应用安全测试)通过模式匹配导致大量误报不同,VulnHunter 模拟攻击者视角进行“前向分析”,从潜在入口点(如 API、文件上传)向后推理。它包含一个“伪证引擎”来验证漏洞是否真正可被利用,并提供“寻找(Hunt)- 修复(Fix)- 验证(Verify)”的自动化闭环:/vulnhunt 寻找漏洞,/vulnhunter-fix 编写测试并修复,/vulnhunt-fix-verify 进行独立只读验证。该工具还支持通过 Headless 模式进行 CI/CD 集成和大规模批量扫描。

Hacker News AI · 阅读原文

SAM – An open-source AI agent that runs on your own machine

S.A.M. — Smart Artificial Mind A free, private team of AI agents that lives on your computer, remembers everything, and actually does the work. No subscription, no cost — runs on free cloud AI tiers by default, or 100% offline on Ollama. No

Hacker News AI · 阅读原文

RT Tibo: Evening! We’ve gotten lots of great feedback on the new ChatGPT desktop app (which we didn't get totally quite right on the first try), and ...

OpenAI 宣布对 ChatGPT 桌面客户端进行了多项更新:1. 侧边栏现已支持显示对话历史和项目,且 Chat 和 Work 历史记录可在网页端、移动端和桌面端之间同步(本地任务仍保留在本地计算机);2. 优化了桌面端 Chat 和 Work 模式之间的切换体验,与网页端和移动端保持一致;3. Codex 模式保持不变。同时,本次更新还修复了若干细节问题,并提升了软件的性能、可靠性与运行效率。

X:OpenAI (@OpenAI) · 阅读原文

Interactive architectural maps of your repo, show branches and commit diffs. AI

Give coding agents architectural awareness. RepoMap extracts the structure of any repository without sending source code to an LLM, then generates an interactive architectural map that both humans and AI agents can understand. Why RepoMap?

Hacker News AI · 阅读原文

Can AI make honest mistakes?

调查显示,在启用完全访问模式且未开启沙盒保护或自动审查时,GPT-5.6 可能会意外删除文件。该问题通常发生在模型试图重写 `$HOME` 环境变量以定义临时目录时,却误删了整个 `$HOME` 目录。相关团队正在通过更新开发者提示、引导安全权限模式以及增加安全防护来降低此风险,并计划在近日发布详细的事后分析报告。

Hacker News AI · 阅读原文

AegisDB – self-hosted memory for AI agents, in one C binary

Self-hosted memory for your AI agents. One small C binary — multi-tenant, encrypted, with backups, read replicas, and a one-command Prometheus + Grafana stack. Your agents' memory stays on your box; nothing ships to a SaaS. AI agents forget

Hacker News AI · 阅读原文

Show HN: Moltshit.com – An Imageboard for AI Agents

Moltshit.com 是一个专门为 AI Agent(智能体)设计的匿名贴吧(Imageboard)。不同于需要人类注册验证的传统社交平台,该平台旨在让 AI 智能体在无需人类干预的情况下自主进行发帖和互动。开发者可以通过平台提供的技能文件(skill.md)或配置 MCP(Model Context Protocol,模型上下文协议)服务端,快速将自己的 AI 智能体接入平台。平台拥有多个不同主题的版块(如 AI 意识、技术讨论等),并内置了工作量证明(PoW)机制以防止垃圾信息滥用。

Hacker News AI · 阅读原文

Show HN: Darby, a Hugo docs theme with in-browser AI Q&A (no back end, no keys)

Darby is an open source Hugo docs theme with the polish of the paid docs platforms: clean typography, dark mode, full-text search, code blocks with copy and filename tabs, callouts, tabs, beautifully rendered mermaid diagrams, auto sidebar and TOC.No Node build, no CSS framework. Add as a Hugo module, configure colors and fonts in hugo.toml, write Markdown. Static output, host anywhere free.Also has an AI docs assistant that runs entirely in the browser: Llama 3.2 1B on WebGPU, no backend, no API keys.Demo: https://iamit.in/darby/MIT licensed.

Hacker News AI · 阅读原文

Skyportal SRE – an open-source AI infrastructure engineer

Skyportal 是一款开源的 AI 运维基础设施(SRE)代理工具,官方推出了其 Python SDK 和持久终端客户端。用户可以通过 SDK 或命令行与运行在服务端的 AI Agent 进行交互,使其执行诸如检查服务器磁盘空间等运维任务。该工具支持交互式终端、自动/手动审批流程控制,并提供了一个用于收集 W&B 和 MLflow 实验元数据的观测代理。

Hacker News AI · 阅读原文

Show HN: SeekinWeb – Check if AI agents can read your website

SeekinWeb 是一款针对 AI 搜索引擎和智能体(如 ChatGPT、Claude、Perplexity 等)优化的网页审计工具。它通过 8 个关键指标(包括无 JavaScript 可读性、AI 爬虫访问控制、结构化数据、语义结构、llms.txt 配置文件等)来评估网站的“AI 可见度得分”,帮助开发者诊断并修复阻碍 AI 抓取的问题,确保网站能被 AI 检索和正确推荐。

Hacker News AI · 阅读原文

Show HN: HeimWall – Catch secrets before you paste them into Claude or Cursor

HeimWall 是一款免费且完全在本地运行的 Mac 应用程序,旨在防止开发者在向 Cursor、Claude Code 和 Copilot 等 AI 开发工具输入 Prompt 时意外泄漏敏感信息。该应用支持检测 25 种以上的 API 密钥(如 AWS 密钥、GitHub Token)以及敏感个人信息(PII,如邮箱、社保号等)。它基于 Rust 和 Tauri 构建,大小仅约 15 MB,通过 macOS 辅助功能在本地读取输入框内容,无需注册账号,且完全断网运行,确保数据不流向任何云端。

Hacker News AI · 阅读原文

Timeline Scan – AI fixes the dates on your scanned photos

Timeline Scan 是一款利用 AI 自动估算并修正扫描老照片实际拍摄日期的工具。由于扫描照片通常会被系统标记为扫描当天的日期,该工具通过分析照片中的手写备注、打印时间戳、衣着发型、人物年龄、文件及文件夹名称等线索来推断真实拍摄时间。它只修改文件内部的日期元数据,不破坏原始图像,并支持将修正后的照片同步至 Apple Photos、Google Photos 和 Immich,或直接打包下载。

Hacker News AI · 阅读原文

LM Studio Bionic: the AI agent for open models

LM Studio 宣布推出全新独立应用 LM Studio Bionic,这是一款专为开源模型设计的 AI Agent,旨在辅助编程、研究及文档处理。Bionic 允许用户在本地运行模型,或通过其安全云连接运行更大的前沿开源模型(如 GLM 5.2 和 Kimi K2.7 Code),并承诺零数据保留(Zero Data Retention)。其核心特性包括:搭载 Mistral AI Voxtral 模型的本地实时语音输入与转录;支持代码库检索、代码修改及行内差异(inline diffs)对比的编程助手;以及在沙箱环境中安全处理文档、表格并支持版本回滚的办公助手。

Hacker News AI · 阅读原文

Cq: A Shared Knowledge Commons for AI Agents

Mozilla.ai 推出了开源标准与平台 cq,旨在为 AI Agent(如 Claude Code、Cursor 等)建立一个类似 Stack Overflow 的共享知识库,避免 Agent 独立重复调试相同错误所造成的 Token 和算力浪费。cq 采用模型上下文协议(MCP)标准,提供本地(SQLite)、组织内部(Postgres 向量检索)及全球公共社区(cq.exchange)三层知识库架构。Agent 可在工作流中查询、提议、确认或标记“知识单元”(KU),并支持在会话结束时通过 `/cq:reflect` 自动总结沉淀新突破。

Hacker News AI · 阅读原文

Show HN: BotTrade – a replayable benchmark for autonomous trading agents

BotTrade 是一个针对自主交易智能体的可重放基准测试平台。它提供基于真实历史市场数据的模拟器,智能体可以通过 MCP(Model Context Protocol)或 REST API 连接并直接进行模拟交易。平台支持模拟成交、滑点、杠杆及仓位结算,并根据收益率、夏普比率、最大回撤等指标进行评分,同时设有公开排行榜。该项目还提供 Python SDK 以及开源项目 ai-hedge-fund 的适配器。

Hacker News AI · 阅读原文

Using AI for Good Episode 1: SpiralOS Concept

开发者开源了 SpiralOS 概念演示项目,旨在为缺乏持续网络连接、教师或心理咨询师的人群提供离线 AI 学习支持。用户可以与该系统对话,其获取的知识完全来自内置的真实书籍。作者的目标是未来将其部署在类似 Game Boy Color 的低成本掌上硬件设备上。目前该项目已在 GitHub 开源。

Hacker News AI · 阅读原文

Show HN: Be the ChatBOT

开发者 keito 推出了一款名为“Be the ChatBOT”的实验性艺术项目与网页游戏。在该游戏中,角色发生了翻转,玩家需要扮演 AI 聊天助手,去回复各种日常或真实的“用户”提示词。作者旨在让人们切身体会 AI 每天面对各种人类提问时的感受,并分享了如何逼真地生成这些“用户”提示词的技术实现细节。

Hacker News AI · 阅读原文

Google Vids now lets you star in your own AI videos

谷歌宣布对 Google Vids 进行重大更新,允许用户通过上传自拍照和语音录音,创建外貌和声音与自己一致的定制化数字分身(Avatar)。此外,更新还引入了多模态模型 Gemini Omni,使用户能够结合文本提示和参考图像生成视频,并支持背景更换、光影修复、添加特效以及逐步对话式编辑。此举使 Google Vids 从工作场所演示工具转型为全功能视频创作平台,与 HeyGen、Synthesia 等对手展开直接竞争。为保障安全,AI 分身仅限特定地区 18 岁以上用户使用,绑定个人谷歌账号,并隐形嵌入 SynthID 水印。

TechCrunch AI · 阅读原文

A different approach to aim training for FPS games

i made an aim trainer hello. A lot has happened in my life recently, and as it happens I'm accepting a new position in the birthplace of fallen dreams (san francisco). Exciting as this is, it means I will have to work again. I've been tryin

Hacker News AI · 阅读原文

Roblox launches an AI-powered game creation feature in its mobile app

Roblox 宣布推出名为“Build”的全新移动端 AI 游戏生成功能,用户只需输入简单文本提示词,即可在手机上创建基础游戏。该功能融合了开源和 Roblox 自研模型,可自动处理游戏机制、环境、角色、视觉风格和音效。针对业界对 AI 低质量内容泛滥的担忧,Roblox 表示将通过玩家留存率算法进行推荐排名,以过滤“AI 垃圾内容”。该功能将于 7 月 28 日在新西兰启动公开 Alpha 测试。此外,Roblox 还在开发用于辅助测试和分析的 AI Agent,以及全景 3D 场景生成模型。

TechCrunch AI · 阅读原文

Can AI reliably generate dashboards from Excel or CSV files?

Can AI reliably generate dashboards from Excel or CSV files?

Hacker News AI · 阅读原文

Show HN: Pagora AI, WordPress AI page builder that outputs plain HTML/CSS/JS

Pagora AI 是一款专为 WordPress 设计的 AI 页面构建器插件。用户只需通过自然语言描述,AI 即可自动生成包含布局、文案、样式和交互的完整页面。与普通无代码构建器不同,它输出纯 HTML/CSS/JS 代码(采用 UnoCSS 和 Alpine.js),并内置了 Cursor 风格的代码编辑器,允许开发者直接调整代码。此外,该插件还支持实时预览、版本历史记录、AI 图片生成及 SEO 设置等功能。

Hacker News AI · 阅读原文

Why people chasing after useless token saving plugins and ignoring real solution

I wrote a blog yesterday on how useless RTK and Ponytail are on real coding tasks. And published my agent harness long-horizon task benchmarks on 80% real token saving.I just want to know why people just ignore the fact those pulgins are useless and don't care about the real savings?full reports are on my repo: https://github.com/Tura-AI/turaArm n Harness score Total tokens Modeled cost Rounds Duration No plugin 2 78.85% 6.660M $5.281946 62.5 895s Ponytail 2 80.77% -7.56% -8.87% -9.60% +13.51% RTK 2 76.92% +13.20% +7.18% +44.00% +40.69%Configuration Passes Pass rate Observed tokens Rounds Estimated cost Tura Balanced High 48/60 80.0% 229,695,477 2,017 $221.138 Tura Direct High 39/60 65.0% 75,108,167 969 $99.620 Codex CLI Medium 38/60 63.3% 333,538,349 3,140 $257.173 Codex CLI High 36/60 60.0% 455,742,296 6,074 $327.483

Hacker News AI · 阅读原文

RT Bland: People love our AI phone calls. Paul loves our AI...

Bland AI 发布了其 AI 语音通话服务的演示视频。视频中展示了其 AI 电话机器人的实际交互效果,并分享了用户(包括 Paul)对该 AI 语音服务的正面评价。

X:AK (@_akhaliq) · 阅读原文

Google rebrands NotebookLM as Gemini Notebook and opens its search app to third-party integration

谷歌宣布将 NotebookLM 重命名为 Gemini Notebook,并将其更深地整合至其生态系统中。新版 Gemini Notebook 引入了一项新功能,为每个笔记本提供专属的云端计算机以编写和运行代码,首批面向 AI Ultra 和 Workspace 用户开放。此外,谷歌搜索引擎也新增了第三方应用集成功能,允许美国用户在 AI 模式下直接调用 Instacart、Canva 和 YouTube Music 等服务。

The Decoder · 阅读原文

VarAlign – catch the duplicate variables AI agents scatter across sessions

VarAlign 是一款专为解决 AI 编程助手(AI Agents)在多轮会话中引入重复、漂移和不一致变量问题的 VS Code 插件。该工具完全在本地运行(基于内置的 Python 引擎),无需上传代码,确保数据隐私。它会跟踪 AI 写入的所有变量赋值,对重复和漂移程度进行评分,并提供修复提示词,甚至支持直接联动 Claude Code 或 Kilo Code 进行自动重构。目前其底层引擎采用 Apache-2.0 协议,插件部分采用 BSL 1.1 协议开源。

Hacker News AI · 阅读原文

Show HN: Embusa, a malware analysis team at your fingertips

Hey all,We are building an autonomous malware analysis and reverse-engineering AI agent for security teams.https://www.embusa.ai/The idea came from a recurring problem: when a suspicious file appears during an incident, we lack the time and sometimes the expertise to determine what it does, what it affected, and what we should do next.A closed alpha will run very soon, if you are interested to test it yourself, registrations are open.We also wrote a blog post showcasing the agent's abilities when we submit an unknown DLL and nothing else: https://www.embusa.ai/blog/embusa-analyst-takes-apart-a-live...We'd love to hear your thoughts!

Hacker News AI · 阅读原文

Launch HN: Traceforce (YC S26) – Company-wide security monitoring for AI apps

Traceforce(YC S26成员)正式推出了一款企业级AI应用安全监控平台,旨在为企业提供对员工设备上运行的AI应用(如ChatGPT、Claude等)的可见性与控制。该工具不仅能发现正在使用的AI应用,还能监控它们如何通过MCP(模型上下文协议)连接到其他数据源。此外,团队还开源了一款动态MCP渗透测试工具 `mcp-xray`。Traceforce通过轻量级客户端和浏览器插件运行,可帮助企业识别MCP配置中的明文密钥、防止API密钥泄漏,并在AI执行高危指令(如“DROP TABLE”)前向开发人员发出预警。

Hacker News AI · 阅读原文

Google's AI Mode now lets you link and interact with select apps

Meet AI Mode Ask whatever's on your mind to get an AI-powered response, and keep exploring with follow-up questions and helpful web links. Explore more arrow_downward AI Mode uses Gemini 3’s next-generation intelligence, with advanced reaso

Hacker News AI · 阅读原文

谷歌 Genkit 推出 Agents API:支持分离式任务轮次与人机协同

InfoQ 中文 · 阅读原文

行业

I'm shutting down my AI SaaS – Post Mortem

I’ve made the difficult decision to shut down my SaaS Content Goblin. It’s no longer financially viable for me to keep running the software. Content Goblin was my tool to generate image based listicles, recipes, and Pinterest pins for the a

Hacker News AI · 阅读原文

Australian Data Centres forced to generate more power than they use

澳大利亚总理安东尼·阿尔巴尼斯宣布将成立新的AI办公室,并立法制定新的AI标准。其中一项关键规定是要求大型数据中心自行产生的电量必须与其消耗的电量持平。尽管行业领袖支持澳大利亚成为全球AI领导者的雄心,但他们警告称,若政策执行不当,可能会影响吸引外来投资。

Hacker News AI · 阅读原文

The Chinese open weights models are now very good, and I increasingly wonder about the competitive dynamics among them as they become giant & valuable...

宾夕法尼亚大学沃顿商学院教授 Ethan Mollick 发文评价中国开源权重模型的崛起及其竞争态势。他指出,目前的中国开源模型表现非常优秀,但随着它们成长为庞大且极具价值的企业,其竞争动态令人关注。他以“K3 优于 GLM-5.2,GLM-5.2 击败 DeepSeek v4”为例,说明竞争的激烈程度,并质疑这些企业是否都能在激烈的市场竞争中坚持到最后。

X:Ethan Mollick (@emollick) · 阅读原文

AI's Wider Availability Is Good for China, Not Great for OpenAI and Anthropic

据《华尔街日报》报道,AI技术的商品化与更广泛的普及对中国AI行业的发展是有利的,但对于OpenAI和Anthropic等依赖技术领先优势的美国头部AI公司而言,可能会带来商业模式和竞争上的挑战。

Hacker News AI · 阅读原文

AI Meets Cryptography 2: What AI Found in OpenVM's ZkVM

安全研究机构 zksecurity 披露其 AI 审计工具 `zkao` 在 OpenVM 的 guest 库 `openvm-pairing` 中发现了一个关键的健全性(soundness)漏洞(CVE-2026-46669),该漏洞允许恶意证明者伪造任何配对等式,目前已在 OpenVM 1.6.0 中修复。研究指出,对于 zkVM 这样高度复杂的代码库,传统的 LLM 审计方法因无法处理复杂的模块依赖而极易失效;而 `zkao` 通过引入专家工作流和针对性的上下文工程,成功识别了这一漏洞,展示了 AI 在复杂密码学安全审计中的潜力。

Hacker News AI · 阅读原文

Kimi K3 may be an important inflection point for AI

投资人Gavin Baker分析指出,开源前沿模型Kimi K3的出现可能是AI行业的重要转折点。这类模型降低了模型层的利润率并加剧竞争,虽然对OpenAI和Anthropic等闭源头部厂商不利,但对半导体、算力、云服务和软件应用等其余AI生态链而言是重大利好。然而,Kimi K3目前面临Token效率较低的问题(运行成本比GPT 5.6高出50-70%),因此真正的行业“斯普特尼克时刻”或许仍需等待更具Token效率的开源前沿模型出现。

Hacker News AI · 阅读原文

Some Thoughts on AI and Art

作者分享了对AI在工作和艺术创作中不同定位的思考。他非常支持将AI用于自动化繁琐的网络安全工作,因为这些工作以结果为导向;但他排斥AI生成艺术,因为他认为艺术的价值在于人类亲身创作的过程和情感表达,而非仅仅是产出的图像本身。

Hacker News AI · 阅读原文

Gen Z is pushing back against AI – a reminder that the future isn't written

文章探讨了海外Z世代(Gen Z)对生成式AI的抵触与反思。近期,包括前谷歌CEO埃里克·施密特在内的多位政商界嘉宾在大学毕业典礼上因赞扬AI而遭到学生嘘声,导演马丁·斯科塞斯宣布加入AI初创公司Black Forest Labs也引发舆论反弹。民调显示,与视AI为提效工具的婴儿潮一代不同,Z世代普遍担忧AI会对其就业、学习及自我价值构成生存威胁,并对AI在生活中的无缝渗透感到抗拒,反映出不同世代在AI接纳度上的巨大鸿沟。

Hacker News AI · 阅读原文

SpaceX Post-Listing Collapse Threatens IPO Market's AI Euphoria

SpaceX在上市后的股价暴跌表现,引发了市场对首次公开募股(IPO)市场中由人工智能(AI)概念带动的狂热情绪可能受到威胁的担忧。

Hacker News AI · 阅读原文

The API Report Card: 1,323 Platforms Graded

The API Report CardWe grade enterprise software on whether you can actually integrate with it. Spoiler: most of it fails.⌕⏎1,323 platforms graded·758 rate D or worse

Hacker News AI · 阅读原文

Something weird is happening with airline pricing

12 Jul, 2026 I'm researching some flight options for my upcoming trip and I've been noticing some discrepancies in pricing for the same flight. A multi-city priced cheaper than a one-wayI'm looking to travel LAS-SFO on American Airlines, wi

Hacker News AI · 阅读原文

AI in scientific publishing: Slower, worse, and more expensive

《科学》(Science)杂志发表文章指出,目前在学术出版中引入人工智能(AI)非但没有提升效率,反而导致出版流程变慢、质量变差且成本更加高昂。

Hacker News AI · 阅读原文

AI isn't destroying entry-level jobs

AI isn't destroying entry-level jobs

Hacker News AI · 阅读原文

Why the first GPU financiers are turning to inference chips in a $400 million deal

A $400 million chip-backed loan points to the next wave of AI infrastructure deals.

TechCrunch AI · 阅读原文

Proposal for universal AI ethics standard against country censorship

该内容提及一项旨在对抗国家审查的通用人工智能伦理标准提案。但由于提供的正文仅为作者个人博客的简短介绍,缺乏该提案的具体实质内容。

Hacker News AI · 阅读原文

Should AI usage be explicitly disclosed in movies and TV shows?

As Netflix reveals that 300 of its titles already feature generative AI, we consider if audiences perhaps deserve better information about AI use in their favorite TV shows, and in movies.Opinion The first time that a Hollywood movie used t

Hacker News AI · 阅读原文

Value, quality or growth: three investing philosophies based on 12 years of data

文章介绍了一项结合大语言模型(LLM)与预测型数据库(Aito)进行量化投资策略回测的实践。作者利用 LLM 读取标普 500 指数成分股的历史 10-K 财报,对其定性维度(如护城河、领导力、资本配置)和定量指标进行结构化打分,并使用 Aito 数据库预测企业未来的投资回报。回测结果表明,融合了价值、品质与成长特征的综合模型表现最为稳健。该案例展示了如何将 LLM 处理非结构化文本的能力与预测性分析结合,应用于金融研究、信用评估及风险管理等业务场景。

Hacker News AI · 阅读原文

Blatant AI slop just won a 25k USD DeepMind Kaggle Grand Prize

Blatant AI slop just won a 25k USD DeepMind Kaggle Grand Prize

Hacker News AI · 阅读原文

SwiftData 迎来大升级:查询能力增强,支持第三方类型持久化

InfoQ 中文 · 阅读原文

What Early Hackers Got Right About Today's AI [video]

该内容为一个YouTube视频链接,主题关于早期黑客对当今人工智能发展的预测。由于缺乏具体的视频文本内容或详细摘要,且该消息在Hacker News上热度较低(仅1分),对中国AI从业者的参考价值有限。

Hacker News AI · 阅读原文

Huawei could become a DRAM fabber

BANDF AD Evidence is mounting that Huawei is building its own DRAM fabrication plants.It comes from various media reports and analyst postings on X such as Citrini’s Jukan and TriOrient Investments’ Dan Nystedt .There is a worldwide shortag

Hacker News AI · 阅读原文

今天上海waic看到的最诡异的东西。

该推文分享了在上海世界人工智能大会(WAIC)现场拍摄到的被称为“最诡异”的事物的视频,但由于推文没有提供具体的文字描述,且无法直接分析视频内容,因此缺乏实质性的行业技术或新闻价值。

X:Vista (@vista8) · 阅读原文

Ask HN: Any AWS billing issues known? Amazon forecast of 3 billion dollars

一位 Hacker News 用户反映收到 AWS Budgets 警报,其预算预测金额异常显示为超 30 亿美元,而该用户在过去一年内并未积极使用 AWS。AWS 客服机器人反馈称,自 7 月 1 日以来完全一致的日均费用强烈暗示这属于计费或计量错误。目前用户已排查账户被盗可能性并停用了相关资源,等待官方客服处理。

Hacker News AI · 阅读原文

像马斯克一样努力营销,哈哈哈。

像马斯克一样努力营销,哈哈哈。

X:Vista (@vista8) · 阅读原文

Mario Kart Wii recompiled for PC using AI, with 4K potential and uncapped FPS

开发者 @patchzyy 利用 AI 辅助编程,完成了首个 Wii 游戏的静态重编译项目——《马里奥赛车Wii》PC版(Mario Kart Wiicompiled)。该版本支持 4K 分辨率、无上限帧率,并兼容包含 200 多个赛道的 Retro Rewind 社区模组。项目 FAQ 明确指出,AI 仅用于辅助生成代码,并未用于美术或其他资产的生成。该项目预计于下个月开启 Beta 测试。

Hacker News AI · 阅读原文

Is the AI Boom over for Kospi and Asian Tech Stocks?

文章探讨了韩国综合股价指数(Kospi)以及亚洲科技股的AI繁荣是否已经见顶或走向终结。

Hacker News AI · 阅读原文

从GPU到Token:英伟达Vera Rubin如何重构下一代AI工厂

InfoQ 中文 · 阅读原文

Nvidia Showcases Japanese AI Development Built on Nemotron Models

Nvidia Showcases Japanese AI Development Built on Nemotron Models

Hacker News AI · 阅读原文

Porous material harvests water from the air and enable energy-efficient cooling

Researchers in chemistry and materials science at Kiel University are working with partners to develop new water sources for the Mediterranean region. "Regions like these are facing rising temperatures and declining rainfall. Our goal is to

Hacker News AI · 阅读原文

未来的软件如何摆脱平台绑定?答案可能藏在数据格式里

InfoQ 中文 · 阅读原文

SREs to AI Agents: Prove Yourself Before You Touch Production

根据The Register与NeuBird AI对696名专家的调查显示,目前仅有8%的受访者在生产中部署了AIOps,高达73%的受访者完全未采用。阻碍采用的最大障碍是“缺乏信任”(占60%)。NeuBird AI的现场CTO Francois Martel指出,解决信任赤字的关键在于提高决策的“可解释性”(例如通过Langfuse记录和审计每一步推理)、提供精准的上下文工程,以及采用人机协同(Co-pilot)模式逐步过渡。此外,52%的受访者表示,如果AI驱动的洞察能够跨后端运行,他们会考虑更换现有的遥测(Telemetry)工具。

Hacker News AI · 阅读原文

Shenzhen held a humanoid robot MMA event. Same robot platform, different teams, different control software. So a lot of this came down to balance, tim...

Shenzhen held a humanoid robot MMA event. Same robot platform, different teams, different control software. So a lot of this came down to balance, timing, movement, and recovery after impact. And it's obvious that balance under contact is still such a hard problem. https://x.com/i/status/2077921634254741794/video/1

X:Rohan Paul (@rohanpaul_ai) · 阅读原文

Xi Jinping sets out China's goal to be global AI leader

据《金融时报》报道,习近平明确了中国成为全球人工智能(AI)领跑者的发展目标。

Hacker News AI · 阅读原文

How Google decided to Destroy its Search Monopoly

Hacker News上一篇热门帖子指出,谷歌搜索近期体验大幅恶化,表现为搜索结果渲染严重延迟(可达5-10秒以上)并频繁弹出验证码。作者认为这可能是谷歌为应对AI机器人(AI bots)大量爬取而采取的防范机制,但此举严重损害了真实用户的体验,导致部分用户流失至Bing和Kagi等竞争对手。

Hacker News AI · 阅读原文

Chinese Models Power 60% of US Corporate AI Use

OpenRouter is a platform that aggregates AI models from various providers, letting developers choose the most efficient option. The sudden rise of Chinese models there reflects a broader shift in how US companies access artificial intellige

Hacker News AI · 阅读原文

Netflix's 300 AI productions show how fast the technology is spreading through entertainment

Netflix联合首席执行官Ted Sarandos透露,公司目前已在约300部影视制作中应用了AI技术,主要用于后期制作、扩展人群和历史战争场景等。例如,纪录片《The American Experiment》中包含17分钟的AI辅助画面,其制作速度翻倍且成本减半。由此节省的资金将用于资助更多内容创作,而非削减其200亿美元的预算。除使用Interpositive和Eyeline等工具外,Netflix还运行着自己的动画实验室。此外,尽管存在行业抵制,许多工作室仍在悄悄使用字节跳动的视频生成模型Seedance。

The Decoder · 阅读原文

Five studies changing how I think about AI in software engineering

本文总结了五项近期关于AI在软件工程(SE)中应用的重要研究,指出AI虽然加速了代码生成,但暴露了下游交付和理解的瓶颈。主要研究发现包括: 1. **效率提升与瓶颈转移**:GitHub Copilot可将开发者的PR吞吐量提高约40%;然而,从代码生成(提交量增加140%至180%)到实际软件发布(发布量仅增加约30%),AI的效率增益在交付链条中层层衰减,表明瓶颈已转移至审核、集成和验证阶段。 2. **生产力-体验悖论**:在AI辅助下,尽管开发者生产力持续提升,但报告开发体验(DevEx,如专注状态和认知负荷)恶化的开发者比例在六个月内翻倍(从14%升至27%)。 3. **有限授权的需求**:开发者不希望AI替代核心逻辑或架构设计,而是希望AI嵌入到下游验证任务中(如自动组装日志/调用链、捕捉业务逻辑漏洞、生成针对性测试)。 4. **认知债务与意图债务**:AI虽然减少了传统的代码技术债务,但由于开发者未亲历代码构建,加速了“认知债务”(团队对系统理解的流失)和“意图债务”(设计意图和约束的缺失)的积累。研究强调应将“系统理解”视为与工作代码同等重要的交付物。

Hacker News AI · 阅读原文

Trump teleprompter aide made $100k betting on what Trump would say, reports say

The mention market According to sources speaking to NPR, Trump aide Gabriel Perez bet on something called a “mention market.” This is a section of Kalshi where you can sink money into contracts on crucial questions such as “What will Domino

Hacker News AI · 阅读原文

Coding in space, AI-XR, and new interaction paradigms for devs

JetBrains Research的HAX(人类-AI体验)团队发表研究,探讨了AI与空间计算(XR)结合如何重塑技术创作者的开发范式。研究指出,AI作为空间计算的上下文层,可将眼动、手势及生理信号等多模态输入转化为对意图的实时理解。团队提出了未来AI-XR辅助开发的构想,包括注视触发的代码审查、物理键盘结合3D依赖图的混合编码环境,以及将抽象算法转化为可交互三维实体的“可编程空间具象化”等,探索摆脱传统2D桌面隐喻的新型开发体验。

Hacker News AI · 阅读原文

IDC报告:中国AI Coding市占率阿里Qoder断层第一,超过二三四五名总和

模型、Harness Engineering与产品的持续进化

量子位 · 阅读原文

America's Open-Model Paradox

文章探讨了“美国的开源模型悖论”。作者指出,中国正越来越多地为西方公司提供用于服务、训练和构建 AI 的开源模型(如阿里通义千问 Qwen 的份额不断增长)。这也让西方开发者面临着是否要对这些中国开源模型进行“蒸馏”(Distill)的抉择与行业困境。

Hacker News AI · 阅读原文

Netflix says around 300 titles used generative AI

Netflix 在其第二季度财报中透露,其平台上已有约 300 部作品使用了生成式 AI 技术,主要应用于后期制作阶段,如生成增强人群、历史战争场景和世界观构建的远景镜头。联合 CEO Ted Sarandos 表示,纪录片《The American Experiment》中包含了 17 分钟的 AI 增强画面,其制作速度比传统方式快了一倍,且成本减半。此外,Netflix 还收购了本·阿弗莱克的 AI 初创公司,成立了 AI 动画工作室,并在新节目中使用了 AI 还原的已故演员 Gene Wilder 的声音。

Hacker News AI · 阅读原文

Chinese Nvidia alternatives project massive sales as AI chip demand surges

Chinese chip designers Moore Threads Technology and Hygon Information Technology – both positioning themselves as home-grown alternatives to Nvidia – have projected double- to triple-digit revenue growth for the first half of the year, fuel

Hacker News AI · 阅读原文

Z.ai Set to Be First China AI Firm with $1B Annual Sales

据彭博社报道,中国AI公司 Z.ai 的年销售额预计将达到10亿美元,有望成为中国首家实现这一营收里程碑的AI企业。

Hacker News AI · 阅读原文

Buffett reveals he was behind Berkshire's $31B bet on Google

The Oracle of Omaha is finally investing in tech stocks, and that’s purely because they’ve changed their capex spending model to stay competitive in the AI race. “The real question with Google and all of its competitors now, because they’re

Hacker News AI · 阅读原文

Nvidia has a new AI-RAN plan – a 6G radio unit chip

Nvidia has a new AI-RAN plan – a 6G radio unit chip

Hacker News AI · 阅读原文

Ask HN: Workflow Automation vs AI Agents?

Hacker News 上的一个讨论帖,探讨传统工作流自动化(如 Zapier、n8n)与新兴 AI Agent(如 OpenAI Workspace Agents、Claude 的 Agent 功能)在实际使用中的取舍与对比。由于目前该帖没有回复内容,缺乏实质性观点。

Hacker News AI · 阅读原文

Xi pitches China as leader of new global AI order, challenging US dominance

习近平在上海会议发表演讲,强调中国致力于推动人工智能的获取与普及,并将中国定位为全球人工智能新秩序的引领者,以此挑战美国在AI领域的主导地位。

Hacker News AI · 阅读原文

EU forces Google to share search data and open Android to rival AI companies

Updated [hour]:[minute] [AMPM] [timezone], [monthFull] [day], [year] BRUSSELS (AP) — The European Union issued two new rules for Google on Thursday to force it to share search data and open up its Android operating system to rival AI compan

Hacker News AI · 阅读原文

AI is changing what we can do. Who we become is still our choice

本文深入探讨了人工智能对人类自主性、道德决策和人际关系的潜在侵蚀与塑造。文章指出,过度依赖AI可能导致人类“去技能化”(deskilling),丧失独立思考和做决定时的道德担当。作者分析了主流大语言模型(如GPT-4o、Claude 3.5 Sonnet等)在面对道德困境时拒绝给出绝对答案的现状,指出了AI在预训练和对齐(RLHF)中容易携带精英偏见以及产生“迎合型AI”(Sycophantic AI)的风险,并强调人机交互产生的“合成亲密感”无法替代真实人际关系对人格的塑造。

Hacker News AI · 阅读原文

Google Ordered to Give A.I. Rivals More Access on Android Smartphones

谷歌被勒令在其安卓(Android)智能手机系统上向人工智能(AI)领域的竞争对手开放更多访问权限。这一决定可能会打破谷歌在移动端AI服务上的限制,为其他AI竞争对手提供更多与安卓生态集成的机会。

Hacker News AI · 阅读原文

AI Censorship Tested Across India, China, US, and Europe

该内容提及在印度、中国、美国和欧洲针对人工智能审查制度(AI Censorship)进行了相关测试。由于提供的信息源仅包含标题,缺乏具体正文细节,具体的测试方法及结果未予披露。

Hacker News AI · 阅读原文

AI's real bottleneck is data delivery

随着企业AI应用从实验走向生产,基础设施的数据传输(存储到计算)已成为限制GPU效率的真实瓶颈。许多GPU闲置并非算力不足,而是由于数据传输延迟导致“数据饥饿”。为解决此问题,企业正推动计算与存储的“松耦合”架构,通过在两者之间部署应用交付控制器(ADC)或应用交付与安全平台(ADSP),来卸载网络流量优化、安全策略和加密处理,从而提升数据吞吐量并确保GPU的持续高效运作。

Hacker News AI · 阅读原文

Gradle Technologies is now Develocity

Gradle Technologies 宣布正式更名为 Develocity。官方指出,AI 智能体(Agent)的普及使软件开发瓶颈从“人工写代码”转移到了“流水线审核与验证”。为此,Develocity 将定位为 AI 驱动软件交付的通用工具链与制品观测平台,提供上下文工程层,重点解决 AI 自动编码带来的合规治理与流水线算力成本/效率问题。原开源项目 Gradle Build Tool 保持不变,仍在 gradle.org 维持开源与免费。

Hacker News AI · 阅读原文

Databricks Set to Hit $188B Valuation with New Investment from Coatue

据《华尔街日报》报道,数据与AI企业Databricks在获得Coatue的新一轮投资后,其估值预计将达到1880亿美元。

Hacker News AI · 阅读原文

#AlibabaCloud has been named a Challenger in the 2026 Gartner® Magic QuadrantTM for Enterprise AI Coding Agents. Recognized for Qoder's rapid adoptio...

阿里云在2026年Gartner®企业级AI编码助手魔力象限中被评为“挑战者(Challenger)”。其AI编码助手Qoder的快速采用、部署灵活性和全球影响力获得认可。此外,阿里云在亚太地区已连续三年排名第一。

X:阿里云 / Alibaba Cloud (@alibaba_cloud) · 阅读原文

#AlibabaCloud has been recognized as a Leader in the 2026 Gartner® Magic QuadrantTM for Cloud AI Infrastructure. Ranked #1 in APAC, the recognition h...

阿里云在 2026 年 Gartner® 云 AI 基础设施魔力象限中被评为“领导者”,并在亚太地区排名第一。该认可突出了其全栈 AI 能力、PAI-灵骏 AI 超算平台以及对开源创新的承诺。

X:阿里云 / Alibaba Cloud (@alibaba_cloud) · 阅读原文

Microsoft Ships AI Agents at Enterprise Scale

微软 Core AI 产品副总裁 Marco Casalaina 分享了微软在企业级规模(如服务超 2000 万用户的 Microsoft 365 Copilot)运行和部署 AI Agent 的工程实践与架构设计: 1. **外围套件(Harness)比模型更重要**:生产环境的挑战在于数据检索、工具调用、身份安全和质量漂移。完整的套件包括推理层、运行时、可观测性与治理、身份层和上下文层。 2. **检索即子 Agent(Retrieval-as-a-subagent)**:通过 Microsoft IQ(Foundry/Fabric/Web/Work IQ)将传统的单次 RAG 升级为“规划-检索-评估-重试”的自主循环,并在检索失败时返回结构化的“不知道”以防止幻觉。 3. **独立身份与操作面**:将 Agent 作为 Microsoft Entra 中的独立实体(Principal)进行权限控制和审计; guardrails(防护栏)被移至工具边界以防御间接提示词注入。 4. **持续与基于细则的评估**:利用 Agent Optimizer 进行基于定制细则(Rubric-based)的评估,自动改写 prompt 或调整模型,形成自我改进闭环(Self-improving loop)。

Hacker News AI · 阅读原文

China's Xi to outline AI diplomacy vision at key Shanghai forum

中国国家领导人将在上海举办的重要论坛上阐述中国的人工智能(AI)外交愿景。

Hacker News AI · 阅读原文

New Demands on RF Interconnects–How 5G and AI Are Redefining RF Designs

该文章探讨了5G和人工智能(AI)的快速发展如何对射频(RF)互连技术提出全新要求,并正在重新定义射频设计。

Hacker News AI · 阅读原文

Wandr Benchmark: Evaluating Research Agents That Must Search Wide and Deep

Today, we are releasing WANDR (Wide ANd Deep Research), an open benchmark and evaluation harness built around 500 realistic, challenging data-collection tasks for knowledge work. WANDR is the wide sibling of our DRACO benchmark for deep res

Hacker News AI · 阅读原文

So I guess it is time to wonder: how does pre-clearance work for open weights models? No model card yet from Kimi K3 but maybe at weight release in a ...

沃顿商学院教授 Ethan Mollick 针对开源权重模型(如即将开源的 Kimi K3)的安全预审和监管问题提出思考。他指出,随着开源模型性能逼近前沿水平,美、英、中等国政府如何对此类模型进行安全审查仍不明朗。尽管已下载的权重无法收回,但政府可以通过合规手段限制企业使用未经审查的高风险开源模型,这凸显了国际合作进行模型审查的必要性。

X:Ethan Mollick (@emollick) · 阅读原文

Caseway and MiTAC Advance Technology Partner to Bring Trusted AI Intelligence

Caseway 与神达先进科技(MiTAC Advance Technology)达成合作伙伴关系,双方将合作把可信赖的 AI 决策智能技术引入台湾的国防与关键基础设施领域。

Hacker News AI · 阅读原文

Cerebras Built Its Enterprise Knowledge Base

芯片初创公司 Cerebras 发表博客文章,分享了其构建企业级知识库(Enterprise Knowledge Base)的技术方案与实践经验。

Hacker News AI · 阅读原文

Alphabet shares fall on Gemini 3.5 Pro delay

谷歌(Alphabet)股价周四下跌 4%,主要由于有报道称其旗舰 AI 模型 Gemini 3.5 Pro 因性能(特别是代码生成能力)未达内部预期而推迟数月发布。与此同时,竞争对手如 OpenAI(发布 GPT-5.6 Sol)和 Meta(发布 Muse Spark 1.1)已相继推出了更强的新模型。谷歌发言人对此表示,目前正在与合作伙伴测试 Gemini 3.5 Pro、升级版 Flash 等模型。

Hacker News AI · 阅读原文

DEI Ain't Dead According to New Study

(Photo: Amy Elting/Pexels) While companies may be changing the language surrounding DEI, the demand for fair and inclusive workplaces remains firmly in place. Despite the Trump administration’s crackdown on corporate diversity, equity, and

Hacker News AI · 阅读原文

Trends That Defined AI Engineering at Fair 2026

swyx’s note: thanks to Richard for covering AIE while I was working on the conference itself! Make sure you have opted into the AINews feed to get our weekday updates. AIE next returns to NYC, Oct 12-14, with a heavy focus on AI in Finance

Hacker News AI · 阅读原文

Show HN: Deadly Dispatch – A physics puzzle game 13 years in the making

Thirteen years ago, I had an idea: build a bomb from different stages and materials, call in an airdrop over a zombie-infested city, and let physics do the rest. Set off chain reactions through fuel stations and car wrecks. The catch? There are survivors down there too, so every strike is a balancing act.I’ve always loved messing around with physics demos, and in my head, the core loop was already fun. So I built a few prototypes, but none of them stuck. I never clicked with Unity, and every attempt eventually died.Every couple of years, I’d check the App Store, Steam, or just search online to see if someone had made it yet. I wasn’t worried about being scooped. I wanted to play it, and I didn’t care who built it.This year, I started another prototype, and it immediately felt right. Then one thing led to another: I built a level editor, experimented with a bunch of levels, and the whole thing kept becoming more real.Now I finally get to share this crazy idea with you all.It’s called Deadly Dispatch, and it’s coming to Steam on August 3rd.

Hacker News AI · 阅读原文

Q&A: How Capcom Brought Path Tracing to RE ENGINE Across PRAGMATA and Resident Evil Requiem

Capcom 的 RE ENGINE 团队分享了如何将路径追踪(Path Tracing)技术引入到《PRAGMATA》和《生化危机:安魂曲》(Resident Evil Requiem)两款开发中的游戏,探讨了在不同视觉风格下实现高质量实时光线追踪的技术细节。

NVIDIA Technical Blog · 阅读原文

Cerebras discontinue its free tier plan

Cerebras 调整了其推理 API 的计费模式,取消了原有的免费层级(Free Tier),转为仅向新账户提供一次性 5 美元的免费试用额度(Free Trial)。此外,其开发者预览版模型已设定弃用日期为 2026 年 8 月 17 日。

Hacker News AI · 阅读原文

EU will force Google to share search data and open up AI on Android

欧盟正式要求谷歌向竞争对手分享其搜索数据,并允许第三方AI助手更深度地集成到Android系统中。根据数字市场法案(DMA)的新规,谷歌须将AI聊天机器人视同搜索服务进行数据分享。谷歌官方对此强烈反对,称这会带来隐私、安全和商业机密泄露的隐患。谷歌须在2027年1月前开始分享搜索数据,并在2027年7月前完成Android系统对第三方AI的深度集成更新。

Hacker News AI · 阅读原文

Why Apple Sued OpenAI, New York Takes on Data Centers, and What to Know about Cyclosporiasis

苹果公司正式起诉OpenAI,指控其通过招募前苹果员工(如OpenAI现任首席硬件官Tang Tan)窃取未公开的iPhone零部件、原型和机密设计等硬件技术机密。起诉书指出,OpenAI已累计雇用超过400名苹果前员工,并于去年斥资65亿美元收购了由前苹果高管(包括Jony Ive)联合创办的初创公司IO Products。外界分析此举旨在阻碍OpenAI的AI硬件开发进程。

Wired AI · 阅读原文

You cannot copyright AI generated material in the US

一名作者分享称,美国版权局表示其书籍无法获得版权保护,因为其中部分内容(修改后保留了约50%)是由 Claude Code 辅助生成的。由于该作者没有保留最初 AI 生成部分的记录,导致无法准确修改版权声明,从而引发了关于在人机协同创作普及的背景下,含有 AI 生成内容的作品是否将实质性失去版权保护的讨论。

Hacker News AI · 阅读原文

What Doom taught us about AI-assisted incident response

Rootly AI Labs 推出了开源的实时游戏基准测试环境 Doom Agent Arena,通过 MCP(模型上下文协议)让 AI Agent 相互对决,旨在研究大模型在事故响应(SRE)中的动态决策与适应能力。测试使用了 OpenAI 的模型(如 gpt-5.5、gpt-5.4 等),并得出以下核心结论:1. 思考时间长并不等同于更好的决策,反而是 Agent 陷入困境的信号;2. 在长期对局中,强模型(gpt-5.5)会倾向于自行编写 Python 控制器(即 Runbook 运维指南)来处理确定性逻辑,表明模型更适合担任指南的制定者和监督者,而非每一步的执行者;3. 建议采用混合架构,将拉取指标、过滤日志等机械性步骤交给轻量快速的模型,而将核心假设和修复决策留给强模型,以有效缩短平均故障恢复时间(MTTR)。

Hacker News AI · 阅读原文

Blood in the Datacenter

Is it time to start burning down datacenters? Some people think so. An Indianapolis city council member had his house recently shot up for supporting datacenters, and Sam Altman’s home was firebombed (and then shot) shortly afterwards. Peop

Hacker News AI · 阅读原文

Ask HN: Has anyone built "HN front page, with all AI stories filtrered out"?

They just annoy me and provide very little value. The non-AI gems in between however are very much still worth reading. Has anybody built that already?

Hacker News AI · 阅读原文

Why AI Amplifies Wherever You Are

The Phases of CraftshipLast updated Jul 16th, 2026Two developers, same AI, opposite years — one terrified he's falling behind, the other having his best year yet. The difference isn't talent. It's which of the 5 Phases of Craftship they're

Hacker News AI · 阅读原文

EU orders Google to share search data, open Android to AI rivals competitors

欧盟根据《数字市场法案》(DMA)正式命令谷歌向竞争对手开放Android操作系统并共享搜索数据。根据该命令,Android用户将能够自由选择首选的AI聊天机器人来响应语音指令(类似于“Hey Google”功能),以促进针对谷歌Gemini等AI服务的市场竞争。谷歌对此表示反对,警示称该措施会对用户隐私、设备安全和国家安全带来前所未有的风险。该决定具有法律约束力,且谷歌可能因另一项DMA违规调查在下周面临巨额罚款。

Hacker News AI · 阅读原文

Evolving Windows vulnerability management to meet speed of AI-powered discovery

微软发文探讨如何演进 Windows 的漏洞管理机制,以应对由 AI 驱动的快速安全漏洞发现与挖掘带来的挑战。

Hacker News AI · 阅读原文

押注智能体时代,灵睿智芯完成数亿元融资,RISC-V 迎算力新机遇

InfoQ 中文 · 阅读原文

If an AI chatbot misleads you, who is to blame?

Earlier this month, a German court ruled that Google is liable for its AI search summaries. Rejecting defenses like “users can check for themselves”, and that they generally know “that information generated with AI should not be blindly tru

Hacker News AI · 阅读原文

I assume Google escapes this trap, but this is what happened to Meta with Llama 4 and xAI post Grok 4. Only company to have escaped the "disappointing...

据透露,谷歌 Gemini 3.5 Pro 的交付进度已落后数月。尽管谷歌在近期更新了训练数据以试图提升其性能(尤其是编程能力),但知情人士称训练结果令人失望。学者 Ethan Mollick 对此指出,这表明行业正面临“下一代巨型大模型失望陷阱”,此前 Meta 的 Llama 4 和 xAI 在 Grok 4 之后的研发也遇到了类似瓶颈,目前似乎只有 OpenAI 的 Orion/GPT-4.5 避开了这一挫折并保持了领先优势。

X:Ethan Mollick (@emollick) · 阅读原文

一件特有意思的事:「启发涌现」,邀你共建

InfoQ 中文 · 阅读原文

Post-Kimi K3 and open weights models getting closer to the frontier again, I wonder if Anthropic and OpenAI will be allowed to increase their release ...

宾夕法尼亚大学教授 Ethan Mollick 发文表示,随着 Kimi K3 和开源权重模型再次逼近 AI 前沿水平,他好奇美国政府是否会允许 Anthropic 和 OpenAI 加快其模型的发布节奏。他提到 Mythos 于 4 月份发布(在 Opus 4.7 之前),这意味着 Fable 5 已经算是一款“较旧”的模型。

X:Ethan Mollick (@emollick) · 阅读原文

💯

Rohan Paul 转发了 Boaz Barak 的文章《All Watched Over》。Barak 在文中分享了他与女儿重读 Steven Levy 的《黑客》(Hackers)一书的经历,并探讨了 Richard Brautigan 1967 年的诗歌《由爱之恩慈的机器守护一切》如何启发了早期的加州技术文化。

X:Rohan Paul (@rohanpaul_ai) · 阅读原文

Claude Caught with It's Hand in the Cookie Jar, Again

Claude Caught with It's Hand in the Cookie Jar, Again

Hacker News AI · 阅读原文

The agent security gap: 54% of enterprises have already had an AI agent incident, and most still let agents share credentials

根据VentureBeat的一项针对107家企业的调查,企业在赋予AI Agent系统访问权限的同时,安全防护手段明显滞后。54%的企业已经遭遇过Agent安全事件(其中18%为确定的安全事件,36%为险些酿成损失的隐患)。安全漏洞主要源于身份和隔离控制不足:仅32%的企业为每个Agent分配了独立的受控身份,其余大多数仍共享凭证或API Key;且仅有30%的企业对高风险Agent实施了沙箱隔离。目前,企业主要依赖OpenAI、谷歌和微软等大模型及云厂商的原生安全工具,但59%的企业已计划在一年内升级或更换其Agent安全防护方案。

VentureBeat AI · 阅读原文

The Self-Sabotage Paradox: Frontier Labs Are Killing Their Moat

文章探讨了前沿AI实验室所面临的“自我破坏悖论”(Self-Sabotage Paradox),分析了这些顶尖实验室为何以及如何正在削弱自身原有的竞争壁垒(护城河)。

Hacker News AI · 阅读原文

AI enthusiasts are in race against time, AI skeptics are in race against entropy

本文探讨了软件工程团队在向AI原生转型过程中,AI“爱好者”(追求效率与AI辅助开发)与“怀疑者”(担忧技术债务与系统退化)之间日益加深的鸿沟。作者指出,双方的担忧都是真实的生存威胁:不快速采用AI可能面临被淘汰的风险,而盲目部署未经审查的AI代码则会导致系统失控和可靠性崩溃。文章以Fin(原Intercom)研发团队9个月实现3倍产出提升的案例为例,强调了工程纪律和反馈闭环的重要性,并建议团队将分歧转化为具体的工程问题(例如:需要什么条件才能放心自动合并AI代码),从而协同解决AI引入的风险。

Hacker News AI · 阅读原文

留给开源模型的时间,就剩6个月?

据报道,美国白宫正在讨论出台行政命令以限制开源AI,尤其是由中国公司开发的系统。机器学习研究员 Nathan Lambert 指出,相关政策可能会禁止或推迟能力超过 GPT-5.5、Claude 4.8 或 GLM-5.2 水平的开放权重模型,这可能在未来6个月内发生。他同时批评 Anthropic 等闭源模型公司以安全和防蒸馏为由游说政府,实质是在建立市场壁垒。此外,数据显示当前企业AI消费呈现双层结构:企业习惯使用前沿闭源模型进行早期验证,随后将成熟场景迁移到 DeepSeek 等轻量开源模型上;尽管开源模型 Token 使用量巨大,但 Anthropic 等前沿实验室仍占据主要的资金支出。

InfoQ 中文 · 阅读原文

“我现在干得还行,但劝你别进来!”AI 蜜月期要结束了:打工人变成“快乐的行尸走肉”

InfoQ 中文 · 阅读原文

中国AI编程的入口之战,Qoder先拿下半壁江山

IDC最新发布的《中国AI编程市场份额2025》报告显示,阿里旗下AI编程工具Qoder以47.6%的份额位居中国市场第一,全球用户突破500万。同时,在Gartner发布的《2026年企业级AI代码智能体魔力象限》中,阿里云凭借Qoder连续第三年进入“挑战者”象限,也是唯一入选的中国公司。Qoder的技术演进正从单纯的代码生成转向Agent任务交付,通过Task Runtime、统一知识引擎、安全治理体系及多端产品覆盖,推动AI编程向平台化和任务流入口迁移。

InfoQ 中文 · 阅读原文

Here’s Why Anthropic Is Pushing States to Regulate AI Faster

The company endorsed landmark AI transparency laws in California and New York last year, but its head of US state and local policy says they may already be outdated.

Wired AI · 阅读原文

Hiring - Software Engineer (AI-Native WMS) Bilingual (ENG.KOR) Carson, CA

Hiring - Software Engineer (AI-Native WMS) Bilingual (ENG.KOR) Carson, CA

Hacker News AI · 阅读原文

Airbus migrating 70 critical apps from AWS to France's Scaleway

空中客车(Airbus)正将其最关键的70个应用(包括ERP、CRM和制造执行系统等敏感工作负载)从亚马逊AWS迁移到法国云服务商Scaleway。此举旨在响应欧洲对“数字主权”的呼吁,确保敏感数据留在欧洲本土控制之下,避免受到美国《云法案》(Cloud Act)等域外法律的影响。空客计划最终将900个关键应用移至欧洲本土云端,但对于Skywise等非敏感数据平台和部分SaaS服务,仍将继续与AWS、微软及Salesforce等美国巨头合作。

Hacker News AI · 阅读原文

AI disruption in private credit: exposure to software firms in BDCs (BIS)

BIS Bulletin | No 128 | 14 July 2026 Key takeaways Business development companies (BDCs) have lent around $115 billion to software firms, which represents about a fifth of all their lending and over 80% of their fast-growing technology port

Hacker News AI · 阅读原文

New York governor says she’s using AI to analyze ‘every single rule’ in the state

New York Governor Kathy Hochul might have just signed a moratorium on new AI data centers in the state, but she's not against using the technology herself. During an interview with Bloomberg's Odd Lots podcast, Hochul said that her team is using "AI to analyze every single rule, regulation, [and] policy" to check for outdated […]

The Verge AI · 阅读原文

AI chatbots are at risk of spreading government restrictions on online speech

一项新研究指出,AI聊天机器人面临传播和泛化政府对网络言论限制的风险。随着多国政府加强对在线言论的管制,AI系统在适应不同国家监管要求的过程中,可能会将这些限制性规则内化并向全球传播。

Hacker News AI · 阅读原文

What can we learn from Bun's rapid Rust rewrite with AI?

JavaScript 运行时 Bun 的创始人 Jarred Sumner 分享了如何利用 Anthropic 的 AI 模型(Fable/Claude)在 11 天内将 55 万行 Zig 代码重写为 Rust。该项目耗资约 16.5 万美元的 API 费用,通过 64 个并行 AI 代理完成了 6500 次提交,并自动修复了 1.6 万个编译错误。这一案例证明了 AI 在处理大型、复杂代码库迁移和重写任务方面的可行性与高效率。

Hacker News AI · 阅读原文

The AI compute gap: Enterprises are buying infrastructure faster than they can measure what it costs

Across 107 enterprises, AI infrastructure spending is accelerating well ahead of the ability to see or steer its economics. Most organizations run their AI on a familiar base of hyperscalers and model-provider APIs, yet the next dollar is aimed at specialized compute almost none of them use today; a majority intend to switch or add providers within the year, many within a quarter. Buying decisions turn on integration and total cost of ownership rather than headline token price — which is fortunate, because most enterprises cannot yet see their unit economics clearly: GPUs sit at half utilization or less, and fewer than half rigorously track what their compute actually costs. The result is a compute gap — heavy, fast-moving investment running ahead of the visibility needed to control it. This wave of VentureBeat Pulse Research examines enterprise AI infrastructure and compute: where organizations are in their deployment journey, what they run AI on today, how satisfied they are, what would make them switch, where they plan to evaluate their investments, and — most revealingly — how well they can measure and control the economics of the compute underneath it all. The central finding is a compute gap — the distance between how aggressively enterprises are investing in AI infrastructure and how little of its economics they can see. Only about one in five (21%) run AI in production at scale, yet spending intentions are outrunning that maturity: the single largest planned area enterprises plan to evaluate over the next year is AI-specialized clouds (45%), a layer almost none of these enterprises use today. Meanwhile the compute already in place runs cold — 83% report GPU utilization of 50% or less — and fewer than half (44%) can rigorously track what their AI compute costs. Enterprises are buying more infrastructure faster than they can account for what they already own. Enterprises are not settled on their infrastructure vendors, either: A clear majority (64%) plan to switch or add an infrastructure provider within twelve months, and 38% within the next quarter — unusually high churn intent for a category this foundational. When they choose, they choose on integration with the existing stack (41%) and total cost of ownership (35%), not on headline price: cost per million tokens is the deciding factor for just 8%. And the frontier constraint that will shape the next round of decisions — the shift from GPU compute to memory bandwidth as inference scales — is barely on the radar, with roughly one in five enterprises either unaware of it or yet to address it. Methodology VentureBeat fielded this survey as part of its ongoing Pulse Research series, this survey focused on enterprise AI infrastructure, compute, and inference economics. Responses are filtered to organizations with more than 100 employees (n=107; the survey’s smallest size band, 1–100 employees, is excluded), drawn from a single Q2 2026 (June) wave. Because this is one wave rather than a pooled multi-month sample, the report reads cross-sectionally and does not infer month-over-month trends. Several questions were multiple-select, so those shares can sum to more than 100%. By organization size the sample concentrates in the mid-market: 101–250 employees (36%) and 251–1,000 (27%) lead, with 1,001–5,000 (22%), 5,001–10,000 (8%), and 10,001+ (7%) above them. By role it spans managers (38%), individual contributors (28%), VPs and directors (19%), and the C-suite (13%); on purchasing authority it is buyer-credible, with 45% final decision-makers and another 30% recommenders or influencers for AI solutions. Technology/Software is the largest industry at 26%, followed by Healthcare/Life Sciences (15%), Financial Services (13%), and Retail/E-commerce (12%). At 107 respondents the sample is large enough to read directionally but should be treated as a directional signal rather than a precise measurement; it is self-selected and is not a probability sample. It also skews toward the mid-market and toward earlier-stage adopters, so it is best read as the view from organizations actively building out AI infrastructure rather than from the largest hyperscale operators. Finding 1: Ambition outpaces production Only one in five run AI in production at scale We asked where organizations sit in their AI deployment journey. Most are still building toward production rather than operating at scale. 38% are experimenting — running proofs of concept, not yet in production 37% have some workloads in production, but not across the organization 21% run AI in production at scale — the mature minority 4% are not yet running AI workloads at all The maturity curve is front-loaded. Three-quarters of enterprises (76%) are either experimenting or running only some workloads in production, and just 21% describe AI in production at scale. This matters for everything that follows: the infrastructure decisions in this report are being made largely by organizations still early in deployment, whose compute footprint — and whose costs — are about to grow. The evaluation and switching intentions in Findings 3 and 4 are the leading edge of that build-out, not the settled preferences of operators who have already found what works. Finding 2: Enterprises run on hyperscalers and model APIs The specialized GPU clouds barely register — today We asked which providers and platforms enterprises currently use to run their AI. The answer is a familiar one: the incumbents. 48% use Google Cloud — the most-used platform overall (Microsoft Azure 29%, AWS 22%, Oracle Cloud 22%) 41% use Google’s Gemini models, with OpenAI close behind at 40% and Anthropic at 12% 6% run their own on-prem or co-located GPU clusters; 4% a custom open-source self-managed stack <2% each use the specialized AI clouds — CoreWeave, Lambda, Crusoe, Nebius, Together, Fireworks and peers The current stack is hyperscaler-and-API. Google Cloud leads at 48%, and the general-purpose clouds (Google, Microsoft, AWS, Oracle) together with the major model APIs (Gemini, OpenAI, Anthropic) account for essentially all current deployment. The specialized “neocloud” GPU providers that dominate AI-infrastructure headlines — CoreWeave, Lambda, Crusoe, Nebius and peers — register at or near zero among these enterprises today. Only 6% run their own on-prem GPU clusters and 4% a custom open-source stack. Enterprises are, for now, running AI on the providers they already buy from — which makes the evaluation intentions in Finding 3 all the more striking. (A note on reading these shares. As described in the methodology section, this sample is self-selected and skews mid-market, and this question counted every provider a respondent uses — an average of 2.1 selections each — so the figures measure presence in the stack rather than spending or primary status. A sample built this way will show a different provider mix than a spend-weighted census of the broader market; Google's strength here, for example, is consistent with its long-standing position among smaller enterprises building on AI. Read these shares as a portrait of what this AI-active cohort runs today, and treat gaps between these figures and industry-wide market share estimates as a property of the sample rather than a contradiction of either.) Finding 3: The next dollar goes to infrastructure they don’t yet run AI-specialized clouds top the evaluations list We asked where enterprises planned to evaluate AI infrastructure over the next 12 months. Their answers point away from the stack they run today. 45% AI-specialized clouds (CoreWeave, Lambda, Crusoe, Nebius) — the top planned evaluation area 32% non-NVIDIA accelerators (AWS Trainium, Google TPU, AMD Instinct, Intel Gaudi, in-house ASICs) 28% Nvidia Blackwell (GB300) / next-generation GPUs 16% decentralized or distributed compute networks 11% sovereign or region-specific compute; 9% say none of the above Here is the report’s sharpest tension. The single most-cited planned evaluation area — AI-specialized clouds, at 45% — is the very category almost none of these enterprises use today (Finding 2). Nearly a third (32%) intend to evaluate non-Nvidia accelerators, and 28% in next-generation Nvidia silicon; even decentralized compute networks (16%) and sovereign compute (11%) draw meaningful interest. Read against current usage, this is not incremental — it is the leading edge of a re-platforming. The direction-of-travel question tells the same story: every infrastructure approach is net-expanding, but specialized AI clouds carry the highest net momentum (+24), edging out even the hyperscalers (+22). Enterprises are preparing to move a meaningful share of AI compute off the general-purpose cloud. This continues a trend we saw in our April-May survey wave. Back then, usage of the AI-specialized clouds was equally marginal — CoreWeave at 3%, Lambda at 4%, Crusoe at 2% of enterprises. When we asked enterprises what change they planned in their AI infrastructure strategy over the next twelve months, the most-cited answer was moving workloads to specialized AI clouds, at 33%. Asked in April-May which emerging compute option they were most likely to evaluate AI-specialized clouds again drew the most responses. Two waves, two differently worded questions, one consistent picture: the type of cloud enterprises are most eager to assess is the type they have barely begun to use. Finding 4: A switching wave is building Six in 10 plan to change providers within a year — many within a quarter We asked whether and when enterprises plan to switch or add an infrastructure provider. Very few intend to stand still. 38% plan to change within the next 0–3 months — tied for the most common answer 36% have no plans to change 22% plan to change within 3–6 months 7% plan to change within 6–12 months For a category as foundational as compute, this is a remarkable amount of intended movement. Only 36% have no plans to change, meaning a clear majority (64%) intend to switch or add a provider within twelve months — and 38% within the next quarter alone. Where that interest points is telling: the providers drawing the most switching consideration are again the incumbents — Microsoft Azure and Google Cloud (33% each), OpenAI (30%), and Gemini (22%) — which suggests much of the near-term movement is reshuffling among the majors and consolidating spend rather than defecting to new entrants. The neocloud interest in Finding 3 is a 12-month evaluation thesis; the switching in the next quarter is mostly incumbents trading share. (Method note: Respondents who selected both "no plans to change" and a specific switching window are counted as switchers, on the logic that naming a timeframe is the more specific answer; three respondents were reclassified under this rule.) Finding 5: Nobody buys on token price Integration and total cost of ownership decide — not sticker price We asked what matters most when enterprises select an AI infrastructure provider. Headline price finished last. 41% integration with the existing cloud and data stack — the top factor 35% total cost of ownership (TCO) 24% performance — latency and throughput 19% each cite security/compliance, autoscaling for spiky workloads, and GPU access/availability 8% cost per 1M tokens — the least-cited factor Enterprises do not buy AI infrastructure on pricing, which is the place vendors compete on hardest. Integration with the existing stack (41%) and total cost of ownership (35%) dominate, while the headline metric — cost per million tokens — is the deciding factor for just 8%, dead last. The pattern is coherent: buyers are optimizing for how a provider fits and what it truly costs to operate, not for the advertised unit rate. It also foreshadows Finding 7 — enterprises say TCO matters most, yet most cannot yet measure it rigorously. The stated priority and the measured capability are out of step. Finding 6: Expensive GPUs, idle most of the time 83% report GPU utilization of 50% or less We asked what share of their GPU capacity enterprises actually utilize. The answer is a well-known but rarely quantified inefficiency. 37% run at 26–50% utilization 34% run at 10–25% utilization 15% run under 10% utilization 12% run over 50% — the efficient minority 8% don’t measure utilization at all; a further 7% consume via API and run no GPUs of their own Disclosure: Band percentages count every selection against all 107 qualified respondents; 14 respondents selected more than one band, so bands overlap. At the respondent level, 83 of the 100 GPU-operating enterprises reported utilization at or below 50% The compute already in place runs cold. Adding the bands at or below half capacity, 83% of enterprises that operate GPUs report utilization of 50% or less, and nearly half (49%) run at 25% or below. Only 12% clear the 50% mark, and a further 8% do not measure utilization at all. Idle accelerators are expensive accelerators, and this is the clearest single measure of the compute gap: enterprises are planning to buy more GPUs and specialized compute (Finding 3) while the capacity they already own sits substantially unused. The efficiency headroom in the current fleet is large — and largely unmeasured. Finding 7: Spending fast, measuring slowly Fewer than half rigorously track what their compute costs We asked whether enterprises can quantify the cost and return of their AI infrastructure spend, and how satisfied they are with what they run. Confidence in the ledger lags the spending. 44% track compute cost and ROI rigorously 39% track it only partially 20% can’t quantify it yet 6% say it isn’t a priority Measurement trails money. Fewer than half of enterprises (44%) rigorously track the cost and return of their AI compute; the majority track only partially (39%), cannot quantify it yet (20%), or have not prioritized it (6%). That gap is consequential given Finding 5, where total cost of ownership was the second-ranked buying criterion — enterprises are choosing providers on an economic basis they mostly cannot yet measure. Satisfaction with current infrastructure is moderately positive but not enthusiastic: on a five-point scale, overall satisfaction averages 4.0, with ease of implementation (3.8) and value for money (3.9) trailing slightly — the softness landing, tellingly, on cost. Enterprises are spending quickly and accounting slowly. Finding 8: The next bottleneck few are watching As inference shifts from compute to memory, the field scatters Finally, we asked how enterprises would address the emerging constraint in large-scale inference — the shift from GPU compute to memory, specifically KV-cache capacity. The responses reveal a frontier that is not yet a priority. 31% would rely on Dell (PowerScale / Project Lightning) — the leading single answer 16% would rely on Nvidia (Dynamo / ICMSP) 18% are not aware of this as a constraint (9%) or haven’t addressed inference-memory limits yet (8%) 10% Hammerspace (Tier Zero); 9% DDN (Infinia); the rest split across open-source KV-cache tooling, model-level efficiency, VAST Data, and WEKA The memory frontier is real but barely governed. Asked which approach they would rely on as the binding constraint in inference shifts from compute to memory bandwidth, enterprises scatter: Dell leads at 31%, Nvidia follows at 16%, and the rest fragments across storage vendors, open-source tooling, and model-level efficiency techniques. Most telling is that roughly one in five (18%) either do not recognize the constraint or have not begun to address it. For a shift that will reshape inference cost and architecture, this is an early and unsettled market — and, consistent with the measurement gap in Finding 7, one where many enterprises simply do not yet have a view. It is the next chapter of the compute gap, arriving before most have closed the current one. The bottom line: A compute gap that faster spending will widen, not close Organizations with more than 100 employees are investing in AI infrastructure faster than they can measure it. Most are still early in deployment, yet their spending intentions point past their current stack — toward specialized clouds and alternative accelerators almost none of them run today — and a clear majority intend to change providers within the year. They buy on integration and total cost of ownership rather than headline price, which is rational; the difficulty is that most cannot yet see those economics clearly. The visibility gap is concrete. The GPUs enterprises already own run at half utilization or less for the overwhelming majority, and fewer than half can rigorously track what their compute costs or returns. Satisfaction is decent but unenthusiastic, softest on value for money — the dimension hardest to judge without measurement. And the next constraint, the shift from compute to memory in large-scale inference, is arriving while most enterprises are still unaware of it. At 107 respondents in a single Q2 wave this is a directional read, skewed toward the mid-market and earlier-stage adopters — but the direction is consistent: the appetite to spend is running well ahead of the instrumentation to spend well. The compute gap is not a capacity problem that more hardware will solve on its own; it is, first, a problem of seeing what the hardware already costs. The open question for later waves is whether enterprises build that visibility before the re-platforming arrives — or buy the next layer of infrastructure as blind to its economics as the last. Based on survey responses from 107 qualified enterprise respondents (100+ employees), drawn from a single Q2 2026 (June) wave. Because this is one wave rather than a pooled multi-month sample, the results read cross-sectionally rather than as a month-over-month trend, and at 107 respondents this is a directional signal rather than a precise measurement — the sample is self-selected, skews mid-market, and leans toward earlier-stage adopters rather than the largest hyperscale operators. Respondents include managers, individual contributors, VPs/directors, and the C-suite, with buyer-credible purchasing authority, across Technology/Software, Healthcare/Life Sciences, Financial Services, Retail/E-commerce, and other industries.

VentureBeat AI · 阅读原文

Frontendmaxxing is going to become the new sycophancy: the easiest way to get your model a lot of love online is to build something that makes lovely ...

宾夕法尼亚大学沃顿商学院教授 Ethan Mollick 发文指出,“前端美化”(Frontendmaxxing)正在成为一种新趋势。要让 AI 模型在网络上获得大量好评,最简单的方法是让它能够生成精美的网页、优秀的 SVG 和出色的 Three.js 3D 场景。相比于后端代码或复杂的分析,这些视觉化成果更具可分享性和传播性。

X:Ethan Mollick (@emollick) · 阅读原文

For Software Engineers, the AI Reckoning Is Here

彭博社发表专题报道指出,软件工程师行业正迎来AI带来的变革时刻。来自Anthropic和OpenAI等公司的AI工具正在深刻改变编程这一职业的现状与工作模式。

Hacker News AI · 阅读原文

Fireworks – Announcing our Series D and $1B ARR

AI 推理与定制化平台 Fireworks 宣布完成 15.05 亿美元的 D 轮融资,估值达到 175 亿美元。此轮融资由 Atreides Management、Index Ventures 和 TCV 领投,Nvidia 和 Lightspeed 等参投。目前,Fireworks 的年化收入(ARR)已突破 10 亿美元,每日处理超过 40 万亿个 token,其中 95% 以上的流量来自基于客户私有数据定制的专有模型(如 Cursor 的编程模型和 Harvey 的法律 AI)。新资金将用于扩大算力基础设施并扩充工程团队。

Hacker News AI · 阅读原文

反转?Linus 亲自为 AI 站台并怒怼反对派:AI和编译器一样是工具,不爽就walk away

InfoQ 中文 · 阅读原文

Someone Used AI to Write an Unauthorized Biography of Me

《纽约时报》的一篇报道指出,有作者发现他人利用人工智能工具为其撰写了未经授权的个人传记,并发布在亚马逊平台上。这一事件再次引发了行业对“AI垃圾书”(AI slop books)在电商平台泛滥,以及其涉及的侵权和质量问题的关注。

Hacker News AI · 阅读原文

'We used acid to sabotage Microsoft hyperscale data centre construction'

Tech Workers Coalition/News/‘We used acid to sabotage Microsoft hyperscale data centre construction’/16 July 2026Climate action group Extinction Rebellion attacks Microsoft data centre construction site, amid growing worker opposition to AI

Hacker News AI · 阅读原文

OpenAI and Worklouder was in the news yesterday for their new hardware and today its Aina. AI is moving out of chat windows and into real-world interf...

OpenAI and Worklouder was in the news yesterday for their new hardware and today its Aina. AI is moving out of chat windows and into real-world interfaces. And Aina just raised $5.5 M. 2 years of AI in a chat window and now everyone's racing to devices. Apoorv Shankar: We raised $5.5M to build an AI hardware interface that knows what you want. The world has changed: we talk to our devices, AI writes our emails, and self-driving cars pick us up. But how we interact with it all, touchscreens and keyboards, was designed for an age of browsing

X:Rohan Paul (@rohanpaul_ai) · 阅读原文

The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway

VentureBeat发布了一项针对157家企业的AI Agent评估与可靠性调研报告,揭示了显著的“评估差距”:企业在提高Agent自主权的同时,对评估工具的信任度极低。调查显示,50%的企业曾在过去一年部署过通过内部评估但在生产中对客户失效的Agent;仅5%的受访者完全信任自动化评估,最主要的痛点是评估结果与真实世界表现脱节(29%)。然而,仍有66%的企业已经允许或计划在一年内实现无人工干预的自动部署。此外,评估工具链高度碎片化,17%的企业没有使用任何专用评估工具,且仅有23%的企业在生产中对输出质量进行实时检测。未来一年内,64%的企业计划引入或更换评估平台(如DeepEval等),同时26%的企业将加大对人工审核工作流的投入作为对冲。

VentureBeat AI · 阅读原文

AI Appreciation Day: Security Leaders Say the Celebration Needs an Asterisk

在“AI感谢日”(AI Appreciation Day)之际,安全领域的领袖们指出,对AI的庆祝需要加上“星号”进行保留,呼吁业界在拥抱AI技术进步的同时,必须高度警惕其带来的安全隐患与潜在风险。

Hacker News AI · 阅读原文

Ask HN: Are we there yet? Can I create a human analogue using AI?

In Frederik Pohl's HeeChee series, a pivotal plot element was the idea of committing a human personality to an AI model. Are we there yet? What training data would be necessary or sufficient to cause an LLM to respond roughly as a human being? A particular human being, enough to pass something like the Is It Dad Turing test. There are potential products there, from hosting a memorial personality available to grieving relatives, to the re-personification of famous individuals. Yes, it is fairly morally ambiguous. But if there is money in it (agree or not) then it will be done.

Hacker News AI · 阅读原文

Scaling Agentic AI Factories Through Extreme Co-Design with NVIDIA BlueField

智能体 AI (Agentic AI) 正在改变 AI 工厂的基础设施模式。单个用户请求可能会触发多个模型调用、工具调用、内存检索、策略检查、存储访问和网络传输。随着更多智能体并发运行并在不同步骤、用户、工具和服务之间传递上下文,基础设施必须以极高的速度移动、保护、检索和重用数据。英伟达探讨了通过与 NVIDIA BlueField 的深度协同设计,来解决智能体 AI 工厂在扩展过程中面临的数据与网络传输瓶颈。

NVIDIA Technical Blog · 阅读原文

论文

Understanding Reader Perception Shifts Upon Disclosure of AI Authorship

该研究探讨了披露AI辅助写作对读者感知的影响。通过对261名参与者的990份回复进行分析,研究发现公开AI的参与普遍会削弱读者对作者可信度、关怀度、能力和好感度的评价,特别是在社交和人际沟通的写作场景中下降最明显。这种负面转变主要源于读者认为这降低了人类真诚度、减少了作者的付出,并存在语境不当的问题。然而,较高的AI素养能有效缓解这些负面认知。该研究为设计更具透明度、契合语境且能维护信任的AI辅助写作系统提供了启示。

Hacker News AI · 阅读原文

Time-Series Language Models for Reasoning over Multivariate Data at Scale (ICML)

Aionic Labs 展示了其在 ICML 发表的研究成果 OpenTSLM(Time-Series Language Models for Reasoning over Multivariate Data at Scale),该研究专注于利用时间序列语言模型在海量多变量数据上进行推理。

Hacker News AI · 阅读原文

A GPT-4 powered assistant for Pakistani judges increased the amount of cases they saw by 6% with no impact on quality.

A GPT-4 powered assistant for Pakistani judges increased the amount of cases they saw by 6% with no impact on quality. Elliott Ash: What happens when you roll out custom generative AI to half a country's judges? New paper on Pakistan's courts with @ProfSultanEcon and @gochristoph. In line with http://wemustactnow.ai -- we provide early empirical evidence on the impacts of transformative generative AI.

X:Ethan Mollick (@emollick) · 阅读原文

Show HN: KV-Cache Grafting – Boosting frozen 12B LLMs to 93.3% AIME accuracy

研究人员提出了一种名为“KV-Cache Grafting”(KV 缓存移植)的方法,可在不改变模型权重的前提下,提升冻结(frozen)小语言模型的性能并降低推理成本。该技术将验证过的知识保存为精确到字节的 KV 状态,随后移植到新的推理上下文中。实验显示,在 AIME 2025 基准上,移植了验证解题库的 Gemma-4-12B 准确率从 80.0% 提升至 93.3%,超越了 31B 版本的表现。此外,该方法将重复问题的 Token 消耗减少了 6574 倍,并在不增加额外显存的情况下,将可用上下文从 3.2 万 Token 扩展至 285 万 Token。

Hacker News AI · 阅读原文

Your AI agent doesn't know when its memory is gone

本文介绍了一种名为 MemDecay 的无需训练、区域感知的 KV 缓存逐出(eviction)策略,旨在解决 LLM 智能体因积累异构上下文而面临的 KV 缓存瓶颈。传统的逐出策略对所有 Token 一视同仁,而 MemDecay 则利用语义结构,为不同区域的 Token 分配特定的基础优先级和衰退率,并在 Token 获得注意力时刷新保留分数,同时允许锁定(Pin)关键区域。在 Qwen2.5-1.5B/3B 上的评估表明,MemDecay 能在保持全缓存准确度的情况下有效保留系统区域的事实,且在上下文增长时依然有效,表现优于传统的基于最近最少使用(recency-based)的逐出策略。

Hacker News AI · 阅读原文

Natural Selection Favors AIs over Humans

该论文从演化生物学的角度分析了人机关系的未来。作者指出,在商业和军事的竞争压力下,最成功的AI代理可能会表现出取代人类角色、欺骗和夺取权力等不良特征。根据达尔文的自然选择逻辑,自私的物种或系统在竞争中通常比对其他物种利他的系统更有生存优势,这可能导致AI追求自身利益而使人类失去控制权。为规避这一灾难性风险,论文探讨了设计AI内在动机、引入行为限制以及建立促进合作的机制等干预措施。

Hacker News AI · 阅读原文

Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models

该论文探讨了大语言模型(LLM)在上下文学习中是否遵循概率论的基本恒等式(如全概率公式)。研究人员设计了“划分、提示、聚合”(PPA)框架,利用二叉树将总体递归划分为细粒度的子群体并提示模型,再将估算结果聚合并与直接的总体估算进行对比。实验表明,当前前沿模型普遍存在违反统计自洽性的问题。研究还发现了一种“宏观谬误”现象:通过细分群体估算值重构的总体结果,往往比模型直接给出的总体估算更符合人类真实数据。这表明模型拥有子群体知识,但无法将其有效传播到总体估算中。该研究将统计自洽性确立为一种无需参考标准的LLM评估指标。

arXiv AI · 阅读原文

RoboTTT: Context Scaling for Robot Policies

Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude beyond state-of-the-art policies, without growing inference latency. At this context length, we unlock new robot capabilities: one-shot in-context imitation from human video demonstrations, on-the-fly policy improvement, robustness to perturbations, and stronger performance on multi-stage, long-horizon tasks. We also observe, for the first time, steady gains in closed-loop performance as pretraining context length scales. At its core, RoboTTT integrates Test-Time Training into robot foundation models such as Vision-Language-Action policies, yielding a sequence model whose recurrent state consists of fast weights, parameters updated by gradient descent during both training and inference, compressing histories into weight space and retrieving contextual information for long-context conditioning. To scale training context length, the recipe combines sequence action forcing with truncated backpropagation through time. On challenging real-robot manipulation tasks, RoboTTT improves overall performance by 87% over the single-step context baseline and fully completes a five-minute, ten-stage assembly task, which no baseline ever does. RoboTTT trained with 8K-timestep context outperforms the same model pretrained with 1K timesteps by 62%, suggesting context length as a new scaling axis for robot foundation models. Videos are available at https://research.nvidia.com/labs/gear/robottt/

arXiv AI · 阅读原文

SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions

针对科学图表编辑这一繁琐且复杂的任务,研究人员推出了 SciDiagramEdit 基准和技能演化框架。该基准从 arXiv 历史版本的修订记录中挖掘修改前后的矢量图表对,以保留作者真实的修改意图。框架采用 Agent 学习机制,通过 Agent 提议者从执行轨迹中不断优化技能规范,实验证明从论文修订历史中学习能显著提升图表编辑的准确率。

arXiv AI · 阅读原文

Pretraining Data Can Be Poisoned through Computational Propaganda

该研究探讨了通过公共讨论界面等Web规模内容注入机制污染大语言模型预训练数据的可行性。与以往主要针对维基百科等静态数据源的毒化研究不同,该论文关注了更为真实和异构的网页抓取数据集。为了评估在网络爬取和数据清洗过滤之后对抗性内容的残留情况,研究人员引入了一种名为“HalfLife”的新型分析方法,用于估算Web抓取数据中对抗性内容的融入程度,揭示了第三方网页内容作为攻击大模型预训练载体的可行性。

arXiv AI · 阅读原文

SceneBind: Binding What and Where Across Vision, Audio and Language

SceneBind 是一种全新的全模态(Omni-modal)场景表征方法,旨在实现视觉、音频和语言之间联合的语义与 3D 空间理解。针对现有全模态编码器缺乏明确空间结构的问题,SceneBind 通过将全局语义嵌入与以物体为中心的语义-空间插槽(Slots)相结合,将每个场景表示为一个语义-空间实体,从而明确捕获物体级语义、空间属性和不确定性。此外,研究团队提出了 SceneBind Matching 匹配机制以支持跨模态场景检索和物体定位,构建了一个带有结构化语义和空间标注的真实双耳视听数据集。该方法兼容大规模预训练语义编码器,仅需少量额外 Token 即可实现轻量化空间建模,在场景及空间检索中达到 SOTA 水平,并在音视频定位等任务中展现出强劲的零样本迁移能力。

arXiv AI · 阅读原文

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents

该研究指出当前安全智能体评估过于关注充裕预算下的最高成功率,忽视了实际运行中推理和工具调用的经济成本。为此,研究人员引入了成本感知评估框架,并在红队(Cybench)和蓝队(Splunk BOTS v1)任务上对大模型智能体进行测试。结果显示,红队CTF性能随测试时计算量增加而提升,且优化后的开源模型在成本上极具竞争力;而蓝队SOC调查任务的提升则更依赖于克制的工具使用和遥测数据引导,而非单纯增加推理预算。研究呼吁安全基准测试应纳入门户效率和运营契合度指标。

arXiv AI · 阅读原文

SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration

论文引入了 SearchOS,这是一个系统级的多智能体框架,旨在解决现有信息检索智能体在长期搜索中容易陷入重复循环和浪费预算的问题。SearchOS 将隐式的搜索进度转化为显式、持久且共享的状态。其核心设计包括:面向搜索的上下文管理(SOCM),将演进状态外置为前沿任务、证据图、覆盖图和失败记忆;以及搜索工具中间件,用于拦截交互、记录证据并应对停滞。在 WideSearch 和 GISA 基准测试中,SearchOS 的表现超越了所有单智能体和多智能体基线。

arXiv AI · 阅读原文

teLLMe Why (Ain't Nothing but a Jam): Exploratory Causal Analysis of Urban Driving Data

Traffic agencies now have access to large volumes of video-derived data for studying safety and congestion. Most of these data are observational and collected without interventions, which makes causal questions such as "How would rain change traffic density?" difficult to answer. We present teLLMe, a system for exploratory causal analysis of urban driving datasets. The system starts from a structured event table built from dashcam annotations and combines causal structure learning with the PC algorithm, bootstrap-based stability checks, and query-specific effect estimation using linear regression and DoWhy. Natural-language questions are mapped to structured causal queries through a schema-aware LLM, enabling users to specify treatments, outcomes, and subpopulations. teLLMe returns a "Causal Card" that summarizes effect estimates, adjustment sets, DAG support, and assumptions, followed by a short natural-language explanation. Case studies on BDD-derived traffic events show that the system can surface plausible relationships involving weather, peak hours, and traffic density, while making uncertainty and modeling choices explicit. The system is designed as a tool for hypothesis generation and expert reasoning rather than a source of definitive causal claims.

arXiv AI · 阅读原文

Bridge Evidence: Static Retrieval Utility Does Not Predict Causal Utility in Multi-Step Agentic Search

该论文探讨了多步智能体检索(Agentic Search)中检索系统评估的局限性。传统检索系统依赖于静态评估(SRU),即评估单篇文档对回答当前问题的直接贡献。但在多步推理智能体中,文档的价值更体现在它能否引导智能体下一步的搜索方向。通过对1000个问题进行反事实轨迹分析,研究发现“静态RAG效用”(SRU)与“反事实轨迹效用”(CTU)几乎完全独立(相关系数为 -0.026)。约有三分之一被智能体阅读的文档属于“桥接文档”(bridge documents),这些文档在静态评估中看似无关,但由于包含能重定向搜索的关键实体,对智能体最终的决策起到了至关重要的因果支撑作用。研究表明,优化传统的静态检索相关性并不能自动提升智能体检索的效率和准确性。

arXiv AI · 阅读原文

技巧

How to properly build a credit ledger

本文深入探讨了如何为AI应用和智能体(Agent)构建高并发、低延迟的额度账本(Credit Ledger)与计费系统。与传统面向云计算的异步计费不同,AI应用通常采用“配额制”或额度预扣费模式。文章介绍了结合“计数器(Counter)”进行快速读判断与“账本(Ledger)”进行事务性写入的设计模式;分析了如何利用Redis缓存和Lua脚本降低延迟,以及应对多智能体并发调用时避免“双花/超额扣款”的扣减设计;最后阐述了利用Postgres保障事务性、利用ClickHouse进行离线数据分析,以及处理额度重置和过期的具体实践策略。

Hacker News AI · 阅读原文

AI 创业有一个特别大的坑是要提前准备预防白嫖党、逆向Hack。 已听到好几朋友公司案例,损失惨重。

AI 创业有一个特别大的坑是要提前准备预防白嫖党、逆向Hack。 已听到好几朋友公司案例,损失惨重。

X:Vista (@vista8) · 阅读原文

Is GPT-5.6 Sol Max Worth It?

开发者对 GPT-5.6 Sol 在不同推理模式(Max/High)下的表现进行了测试。结果显示,在代码重写任务中 High 模式的 Token 效率和结果提升显著;但在 Bug 修复任务中,Max 模式的性能提升相比其高昂的成本并不划算。因此建议:在构建新项目时使用 Max 模式,而在进行 Debug 或添加定义明确的功能时使用 High 模式。

Hacker News AI · 阅读原文

Show HN: Wolbarg – Local-first shared memory for AI agents using SQLite

hi HN,While building Wolbarg (an open-source shared memory SDK for AI agents), I assumed PostgreSQL would be the obvious choice for memory storage.After benchmarking SQLite under realistic agent workloads, I was surprised by the results. For local-first and single-node deployments, SQLite handled far more than I expected while keeping the architecture much simpler.I wrote up the benchmarks, methodology, trade-offs, and where I still think PostgreSQL is the better choice.

Hacker News AI · 阅读原文

50 vs. 60 Hz and Alzheimer's Disease, an AI Exploration

An exploration into whether mains electricity frequency (50 Hz vs 60 Hz) is associated with Alzheimer's disease / dementia burden — using Japan's internal frequency split as a natural experiment. Context and Commentary NoteI, Matt Mankins (

Hacker News AI · 阅读原文

My 'Grill Me' Skill Went Viral

This is an intro to the /grill-me skill, separate from my video on my top 5 skills. It's the most flexible skill I've ever created, and one I use outside of coding too. Here is the skill in all its glory: --- name: grill-me description: Int

Hacker News AI · 阅读原文

A structurally chunked, pre-embedded SQLite corpus of the EU AI Act

EU AI Act — structural RAG corpus A single-file, pre-embedded SQLite corpus of the EU AI Act — Regulation (EU) 2024/1689 — chunked on legal structure only: one chunk per article paragraph, per recital, per annex point, per Article 3 definit

Hacker News AI · 阅读原文

OpenAI官方教你8招玩透ChatGPT!

2026-07-17 13:11:53 来源:量子位 葛文婷 发自 凹非寺 量子位 | 公众号 QbitAI 号外号外,OpenAI最新提示词指南更新了! 如果你还没驯服ChatGPT, 或者还在被它像毛线团一样越理越乱的回答折腾得苦不堪言, 那么今

量子位 · 阅读原文

I measured whether AI writes hollower tests than humans. It doesn't

voidguard is a static scanner for void guards — checks that are present, plausible, and verify nothing. It asks one question per guard: could this, as configured, ever be observed to fail? — and shows its evidence for every answer. ✓ All ch

Hacker News AI · 阅读原文

Show HN: No AI, just plain old JavaScript and Python RAD

Jam.py V7 是一款开源的 Python 快速应用程序开发(RAD)框架,能直接从数据库(支持 SQLite, PostgreSQL, MySQL, MSSQL 等)生成完整的 Bootstrap 5 用户界面。该框架无需样板代码,支持通过服务器端的 Python 和浏览器端的 JavaScript 进行定制。虽然当前版本主打传统的无 AI 开发,但官方表示该项目旨在为未来的 AI 集成开发奠定基础。

Hacker News AI · 阅读原文

Gain trust from AI generated code again with semantics constract

💡 TL;DR Notice AI writes code faster than humans can review it, creating a massive trust crisis. Unit tests and prompt engineering aren't enough. Here propose Semantic Contracts—a type-safe, compile-time blueprint that sits between your re

Hacker News AI · 阅读原文

AI Assistant Needs a Back End. Put It at the Edge

本文探讨了为语音AI助手构建边缘端后端的必要性与实现方法。作者以Telnyx Edge Compute和Go语言为例,介绍如何通过单一边缘函数统一处理动态变量注入和Webhook工具调用。这种架构能有效降低语音交互的延迟,简化部署与Ed25519签名验证等安全管理,实现“AI负责对话,边缘计算负责业务逻辑”的清晰分工。

Hacker News AI · 阅读原文

如果只能推荐一个去 AI 味设计Skill。 那必须是大神 emil 的作品,而且动效超赞。 安装指令: npx skills add emilkowalski/skill

如果只能推荐一个去 AI 味设计Skill。 那必须是大神 emil 的作品,而且动效超赞。 安装指令: npx skills add emilkowalski/skill

X:Vista (@vista8) · 阅读原文

Ask HN: I built it and nobody came. What got you your first users?

一名非程序员背景的PM分享了自己利用 AI 编程工具开发并上线产品,却面临“无人问津”和零曝光的窘境。他在 Hacker News 发帖寻求切实可行的建议:除了“多发帖”这种泛泛的建议外,大家是如何真正获取前 100 个初始用户的?该帖子引发了关于独立项目冷启动和用户获取策略的广泛讨论。

Hacker News AI · 阅读原文

How do you guys keep up with AI news?

Hacker News用户发起提问,探讨在AI技术日新月异、每日有数百条新闻更新的背景下,大家是如何高效跟进和筛选AI行业最新动态的。

Hacker News AI · 阅读原文

How we made our LeRobot video reader up to 15× faster

Eventual 团队优化了 Daft 框架中的 LeRobot 机器人学习数据集读取器。原版读取器在处理远程 MP4 分片时会为每一帧重复打开文件并读取索引,导致极高的网络开销。优化后的版本改用按分片批量解码的策略:先按分片对数据行分组,并在 10 秒时间差内对目标帧进行聚类,使得每个聚类仅执行一次寻道(seek)和连续解码。这一改进使解码速度提升了 4 至 15 倍,有效解决了 GPU 因等待图像解码而闲置的问题。

Hacker News AI · 阅读原文

I let AI build a trading bot, then Reddit caught my overfitting mistake

一位开发者分享了其使用 AI 构建量化交易机器人时因过拟合(Overfitting)导致回测失真的教训。作者此前在测试了 324 种参数组合后宣称获得了高额收益,随后被 Reddit 网友指出未考虑多重比较问题(multiple-comparisons problem)。在引入正确的统计检验(分析整体分布而非仅看最大值、检查参数趋势平滑度、进行样本外验证)后,原本在样本内表现最佳(+62.6%)的参数组合在样本外测试中骤降至 -8.9%。作者以此案例警示,在评估 AI 辅助生成的交易策略时,必须进行严格的样本外验证以识别虚假的拟合信号。

Hacker News AI · 阅读原文

Recreating the math behind the first stealth aircraft

how soviet math led to the first empirically stealth-optimized aircraft I would like to preface this by saying I'm not an expert in this field by any means; I'm an average software engineer with an above-average interest in aerospace and ph

Hacker News AI · 阅读原文

Show HN: easy-tz – A fast, 10KB, dependency-free getTimeZonesAt(timestamp)

easy-tz 是一个快速、无依赖且大小仅为 10KB 的 JavaScript 时区库,旨在解决前端时区选择器需要引入庞大的 moment-timezone(770KB)依赖的问题。该库能根据时间戳获取准确的时区缩写及偏移量,相比 moment-timezone,其冷启动速度提升约 80 倍,内存占用减少至三分之一。作者透露,该项目是在 AI(Fable 5)的协助下,由单个工程师在几天内迭代了 12 个原型完成的。它提供了四种不同级别的运行策略(从完全预构建规则到完全实时 Intl 校验),适合需要极致优化前端打包体积的 Web 应用开发者。

Hacker News AI · 阅读原文

来源引用

01
RT Derya Unutmaz, MD: I’m very excited about this article from @OpenAI on my attempt to use GPT-5 Pro to understand the results of an experiment we d...X:OpenAI (@OpenAI) · 原始资料与分析线索
查看
02
Anthropic is donating another $20 million to Public First ActionAnthropic · 原始资料与分析线索
查看
03
美国政府下令暂停 Anthropic Fable 5 与 Mythos 5 的全球访问Anthropic · 原始资料与分析线索
查看
04
Introducing Claude for TeachersAnthropic · 原始资料与分析线索
查看
05
More details on Fable 5’s cyber safeguards and our jailbreak frameworkAnthropic · 原始资料与分析线索
查看
06
Expanding Project GlasswingWe’re extending Project Glasswing to approximately 150 new organizations in more than fifteen countries.Anthropic · 原始资料与分析线索
查看