- 2026.10.01 | RIDE外推教师残差;UniEvo-VL在线自蒸馏多模态

 - 2026.10.01 | RIDE外推教师残差;UniEvo-VL在线自蒸馏多模态

Accedi o registrati per recensire o aggiungere questo elemento alla tua collezione.

Tipo d'elemento non supportato.
UUID: 2FexqjumrzkQYR7DMDcNEZ
Class: podcastepisode
Category: podcast

/ 10

0 valutazioni

Non ci sono abbastanza valutazioni
Parte della Programma Podcast: HuggingFace 每日AI论文速递
Sinossi

【赞助商】
OpenClaw快报
每天五分钟,听听 OpenClaw 快报,带你了解最新动态和业内讨论
传送门 https://www.xiaoyuzhoufm.com/podcast/6a1732a2dffa135d0ab5ef43

【目录】
本期的 15 篇论文如下:

[00:27] 🧭 The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation(教师是方向,而非终点:在在线策略蒸馏中外推强化学习诱导的表征残差)
[01:11] 🪞 UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement(UniEvo-VL:面向多模态模型自我改进的在线策略自蒸馏训练方案)
[01:54] 🕵 False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents(虚假前沿:诊断与缓解自演化搜索智能体中的共同作弊)
[02:42] 🤖 AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks(AREX-2:通过长时程反思任务推进自我改进智能体)
[03:29] 🖥 Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents(Mid-Harness:在模型与执行框架之间扩展终端智能体的动作)
[04:09] 🧬 EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery(EvoDuet:面向科学发现的网络搜索与任务求解双层协同演化)
[05:01] 🕵 WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents(WorldAuditBench:使用多模态智能体进行交互式3D世界审计)
[05:47] 🛠 Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI(测试时 AI4AI 中面向智能体执行框架设计的元技能学习)
[06:26] 🧠 EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making(EVOKE:激发智能体中的世界知识以实现可迁移决策)
[07:12] 🎮 RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement(RSIGame:具备递归自我改进能力的自主智能体游戏开发)
[07:53] 📉 More Choices, Fewer Decisions: Ordinal-Scale Bias in JEV-like Direct-Decision Models(更多选择,更少决策:类JEV直接决策模型中的序数尺度偏差)
[08:45] 🧠 Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering(Imagine3D-LLM:教会多模态大语言模型在回答前想象3D场景)
[09:31] 🛠 Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training(智能体错误数据集:面向失败分析与错误感知后训练,规模化构建5万条错误—诊断配对)
[10:21] 🏮 LANTERN: Illuminating Hidden Mathematical Knowledge in Language Models(LANTERN:照亮语言模型中的隐藏数学知识)
[11:05] 🖼 It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them(并非图像所示:无关上下文会扰乱 VLM 评判模型却不为其提供信息)

【关注我们】
您还可以在以下平台找到我们,获得播客内容以外更多信息
小红书: AI速递
在小宇宙查看该单集文稿

Commenti
Recensioni
Notes