alphaXiv

Explore

Researchers

Sign In

MCP Server

Autoresearch

Browser Extension

BlogSend Feedback?

Follow the latest research

alphaXiv connects papers, researchers, and organizations, grounding its answers in the underlying work.

Sign up

Score Centering Stabilizes Off-policy Reinforcement Learning

Together AI
Martin MarekMartin MarekMax RyabininMax Ryabinin

Subtracting the sampler’s expected token scores cancels mismatch-driven drift, stabilizing language-model reinforcement learning under severe quantization and stale rollouts.

17 Sept 2026
632views

Reinforcement Learning for Real-Time Vision-Language-Action Policies

Stanford
Perry DongPerry DongKuo-Han HungKuo-Han HungChelsea FinnChelsea Finn

A lightweight policy can correct delayed vision-language-action commands using current observations, enabling reinforcement-learning adaptation for dynamic real-world manipulation.

16 Sept 2026
3kviews

JEPA-Anything: Learning Predictive Models across Different Worlds

Taoyong CuiZhongyao WangWanli OuyangWanli Ouyang

A shared factorized predictive core supports reusable latent state representations across vision, biology, clinical forecasting, control, molecular dynamics, and physical systems.

17 Sept 2026
120views

Researchers to follow

View all
Yann LeCun

Yann LeCun

Executive Chairman

AMI - Advanced Machine Intelligence, Jacob T. Schwartz Professor, CS @ New York University

Alex L. Zhang

Alex L. Zhang

CS PhD Student

Massachusetts Institute of Technology, Research Fellow @ Prime Intellect

Ion Stoica

Ion Stoica

Co-Founder & Executive Chairman

Anyscale, Co-Founder & Executive Chairman @ Databricks, Professor, CS @ UC Berkeley

Kaiming He

Kaiming He

Distinguished Scientist

Google DeepMind, Associate Professor, EECS @ MIT

Ilya Sutskever

Ilya Sutskever

CEO and Co-Founder

Safe Superintelligence Inc

Chelsea Finn

Chelsea Finn

Co-Founder

Physical Intelligence, Assistant Professor, CS and EE @ Stanford University

Andrej Karpathy

Andrej Karpathy

Researcher

Anthropic

Li Fei-Fei

Li Fei-Fei

Co-Founder and CEO

World Labs, Founding Co-Director @ Stanford HAI, Sequoia Professor, CS @ Stanford University

Are you a researcher? Find your profile

SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness

NVIDIANTU
Haozhe LiuTian YeSong HanSong Han

As coding agents move from supervised code completion to unattended, around-the-clock exploration, their work expands from isolated predictions into long trajectories of reasoning, tool use, and feedback. Token efficiency therefore becomes important for scaling recursive self-improvement. We take an RSI-inspired approach at the harness layer, scaling auto-research loops across increasingly numerous and diverse environments for harness rollouts. At this scale, the process yields reusable improvements that transfer beyond their development setting, moving automated harness discovery toward production-level outcomes. Four mechanisms survive selection and form SoL-Pi, spanning action execution, context compaction, observation handling, and delegated reading. On the 51-task EdgeBench evaluation, SoL-Pi achieves performance comparable to Pi across GPT-5.6 Sol and Opus 5 while reducing recorded token traffic by 44.7-49.0% and API cost by about one third. In other words, estimated hourly savings are $8.75-$13.50 relative to native Codex and Claude Code harnesses, and $4.36-$5.71 relative to Pi.

17 Sept 2026
2k

In-Context Robot Learning with VLM Agents

Shanghai Innovation InstituteHUST
Dongzhou ChengTaoran YiJiaqi WangJiaqi Wang

Vision-language agents can adapt robot behavior from human videos, goal images, and interaction history without updating task-specific parameters.

16 Sept 2026
527views154

What Does Privileged Information Add to On-Policy Self-Distillation?

NUS
XiuYu ZhangWei ChowTat-Seng ChuaTat-Seng Chua

Answer-matched experiments show that most reasoning gains come from distillation itself, while reference benefits vary with the model and student training trajectory.

17 Sept 2026

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

Deepseek
DeepSeek-AIAnyi XuDamai DaiDamai Dai

Cross-layer KV reuse, FP4 caching, and bounded replay make million-token multimodal agents more practical by sharply reducing context-memory requirements.

17 Sept 2026
117views

RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

ZJUAlibaba Group
Yan YuZhengxi LuZhengxi LuYongliang ShenYongliang Shen

Training agents with privileged guidance early, then removing it when alignment stalls, improves reward-driven learning without requiring skills at inference.

17 Sept 2026

DexTouch-WM: Learning Action-Conditioned Tactile World Models from Human Touch for Dexterous Robot Manipulation

PKUHKU
Yan QinYue ChenYue ChenPing LuoPing Luo

Aligned human touch data improves dexterous robots’ predictions of future visual and contact dynamics without requiring additional robot interaction.

17 Sept 2026

An Empirical Study of Harness Design for Coding Agents

UMass AmherstEmory University
Run-Ze FanZihao ZhangSimin Ma

Coding-agent harnesses should be tailored to model capability, task type, and context budget rather than adopted as universal defaults.

17 Sept 2026
137views

How Model Growth, Recursion, and Boundary Operators Influence Scaling Exponents

NYU
Zixi ChenAkshay VegesnaAndrew Gordon WilsonAndrew Gordon Wilson

Architectural depth growth and boundary normalization can improve language-model scaling exponents, making compute-efficiency gains compound as pretraining budgets increase.

17 Sept 2026
225views

Researchers to follow

View all
John Schulman

John Schulman

Co-Founder and Chief Scientist

Thinking Machines

Geoffrey Hinton

Geoffrey Hinton

Emeritus Professor, CS

University of Toronto

Andrew Ng

Andrew Ng

Managing Partner

AI Aspire, Managing General Partner @ AI Fund, Founder @ DeepLearning.AI, Adjunct Professor, CS @ Stanford University, Chairman and Co-Founder @ Coursera

Sergey Levine

Sergey Levine

Co-Founder

Physical Intelligence, Associate Professor, EECS @ UC Berkeley

Demis Hassabis

Demis Hassabis

Chair

Google DeepMind, Chief Scientist @ Alphabet, Founder & CEO @ Isomorphic Labs

Christopher D Manning

Christopher D Manning

General Partner

AIX Ventures, Senior Fellow, HAI @ Stanford University

Yoshua Bengio

Yoshua Bengio

President and Scientific Director

LawZero, Founder and Scientific Advisor @ Mila - Quebec Artificial Intelligence Institute, Canada CIFAR AI Chair @ CIFAR, Full Professor, CS @ Université de Montréal

Yejin Choi

Yejin Choi

The Dieter Schwarz Foundation Professor, CS & Senior Fellow, HAI

Stanford University, Distinguished Scientist, Language and Cognition Research @ NVIDIA

TabPFN-3.5: Technical Report

Benjamin JägerNick EricksonNick EricksonYann LeCunYann LeCun

A tabular foundation model now handles temporal, grouped, text-rich, multimodal, and wide datasets while improving predictive distributions and relational learning.

15 Sept 2026
166views

SplashSplat: Reconstructing Splashing Liquids from Real-World Multi-View Videos

EPFLGoogle
Peiyu LiuDingxi ZhangMarc PollefeysMarc Pollefeys

A calibrated multi-view benchmark and observation-guided representation make real splashes reconstructable for novel-view rendering, motion interpolation, and appearance editing.

17 Sept 2026

OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher

ETHZNVIDIA Research
Damiano Da ColMaximilian IglMarco PavoneMarco Pavone

A render-free reinforcement-learning teacher can improve camera-based driving policies through on-policy supervision, avoiding the exploration burden of rendered reinforcement learning.

17 Sept 2026

Stable and Unstable Singularities in Navier-Stokes

Orson Mengara

This document provides an extended and rigorous framework dedicated to the geometric analysis of the 3D incompressible Navier-Stokes equations [4]. We comprehensively develop geometric proofs related to decay estimates, blow-up profiles, and the foundational partial regularity theory of Caffarelli, Kohn, and Nirenberg (CKN). We examine in detail the Hausdorff dimension of potential singular sets, local monotonicity inequalities, and the spectral analysis of linearized operators to classify singularities into stable and unstable manifolds.

18 Sept 2026

ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments

Hejia GengZesen HuangLing YangLing Yang

Expert-verified scientific code environments let agents learn from executable simulations, graded feedback, and reusable tasks spanning diverse research software.

16 Sept 2026
258views4

Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening

SJTUShanghai AI Lab
Yizhuo LiJianhao YanYu ChengYu Cheng

Sparse, well-separated critic supervision helps PPO track changing state values within long reasoning responses and learn stronger policies.

16 Sept 2026
186views

PointZero: 3D Point Track Completion for Learning Transferable 3D Dynamics

CMUColumbia
Bardienus P. DuisterhofBardienus P. DuisterhofKaifeng ZhangKaifeng ZhangYunzhu LiYunzhu Li

Predicting dense 3D motion from sparse tracks enables robot-free dynamics pretraining that transfers to unseen objects and manipulation tasks.

16 Sept 2026
116views

Agora: Git as Shared Memory for Collective AutoResearch

NVIDIA
Yifan ZhangYifan ZhangYunheng ZouJan KautzJan Kautz

A Git-backed research graph lets independent coding agents preserve, verify, and build on discoveries instead of restarting experiments from scratch.

16 Sept 2026
162views

Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations

Goodfire
Leon BergenLeon BergenUsha BhallaUsha BhallaAtticus GeigerAtticus Geiger

Simple activation probes can detect reward hacking during long agent rollouts, predict later exploits, and reveal cheating behaviors that language-model monitors miss.

16 Sept 2026
127views
There are no more papers matching your filters at the moment.
Sign in

Assistant