Follow the latest research
alphaXiv connects papers, researchers, and organizations, grounding its answers in the underlying work.
Researchers to follow
View allAre you a researcher? Find your profile
As coding agents move from supervised code completion to unattended, around-the-clock exploration, their work expands from isolated predictions into long trajectories of reasoning, tool use, and feedback. Token efficiency therefore becomes important for scaling recursive self-improvement. We take an RSI-inspired approach at the harness layer, scaling auto-research loops across increasingly numerous and diverse environments for harness rollouts. At this scale, the process yields reusable improvements that transfer beyond their development setting, moving automated harness discovery toward production-level outcomes. Four mechanisms survive selection and form SoL-Pi, spanning action execution, context compaction, observation handling, and delegated reading. On the 51-task EdgeBench evaluation, SoL-Pi achieves performance comparable to Pi across GPT-5.6 Sol and Opus 5 while reducing recorded token traffic by 44.7-49.0% and API cost by about one third. In other words, estimated hourly savings are $8.75-$13.50 relative to native Codex and Claude Code harnesses, and $4.36-$5.71 relative to Pi.

Vision-language agents can adapt robot behavior from human videos, goal images, and interaction history without updating task-specific parameters.

Researchers to follow
View allAndrew Ng
Managing Partner
AI Aspire, Managing General Partner @ AI Fund, Founder @ DeepLearning.AI, Adjunct Professor, CS @ Stanford University, Chairman and Co-Founder @ Coursera
This document provides an extended and rigorous framework dedicated to the geometric analysis of the 3D incompressible Navier-Stokes equations [4]. We comprehensively develop geometric proofs related to decay estimates, blow-up profiles, and the foundational partial regularity theory of Caffarelli, Kohn, and Nirenberg (CKN). We examine in detail the Hausdorff dimension of potential singular sets, local monotonicity inequalities, and the spectral analysis of linearized operators to classify singularities into stable and unstable manifolds.

Expert-verified scientific code environments let agents learn from executable simulations, graded feedback, and reusable tasks spanning diverse research software.
































