PromptAI

Research — AI news & research

Benchmarks, datasets, training techniques and the papers pushing the state of the art. Updated continuously from 20+ curated sources.

AgentsVisionRoboticsPolicyResearchToolsLLMs
arxiv · Computer Vision

Point2Part: Unified 3D Partitioning from Point Prompts

Existing 3D part decomposition methods do not necessarily partition the original shape into non-overlapping parts that collectively cover the entire shape, allowing overlaps or gaps that hinder downstream part-level app…

arxiv · Robotics

Skill-Space Shooting for Autonomous Robot Policy Improvement

Robots deployed in the physical world must be able to improve beyond their initial training as they encounter new situations and failures. For this improvement to scale across tasks, it must make effective use of experi…

arxiv · Computer Vision

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evi…

arxiv · Machine Learning

Breakdown of Local Denoising as Semantic Speciation

The dynamics of generative models exhibit two apparently distinct temporal windows: a speciation window, in which a sample commits to a semantic class, and a nonlocality window, in which local context windows become ins…

arxiv · Robotics

Counterfactual Video Generation Enables Scalable Humanoid Loco-Manipulation

Teaching humanoids loco-manipulation skills, such as carrying diverse objects, via visual imitation is a promising path toward generalist robots. However, collecting diverse, high-quality interaction videos, such as cli…

arxiv · Computer Vision

Adversarial Training for Pixel Diffusion

Pixel diffusion models generate RGB images directly, avoiding the bottleneck of an autoencoder, yet their outputs still systematically underrepresent fine-scale natural-image statistics. We show that adversarial learnin…

arxiv · NLP / LLM

STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization

Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low…

arxiv · Machine Learning

LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization

Recent LLMs increasingly adopt hybrid designs that replace standard attention with linear attention, such as Gated DeltaNet (GDN) and Kimi Delta Attention (KDA). Although they compress the context into a fixed-size recu…

arxiv · Computer Vision

Cropland PAtteRNS: Parallel Dimensional Attention Networks and Attention to Dataset Disparity for Crop Segmentation in Satellite Imagery Time Series Data

The landscape of satellite imagery time series datasets and boundary-pushing architectures for cropland segmentation has never been richer. However, in this gold rush, important truths are being missed on both fronts, a…

arxiv · Computer Vision

Rethinking Representations for World-Action Modeling

World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and prediction. We study the design of this space through controlled compari…

arxiv · Machine Learning

A Spectral Theory of Distortion in LLM Graph Reconstruction: Sharp Bounds and Empirical Characterization

Evaluations of graph reconstruction by language models typically report a single aggregate distance between the original and the reconstructed graph. We prove that for the Wasserstein distance between Laplacian spectra…

arxiv · NLP / LLM

EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation

Emotion-conditioned text-to-speech (TTS) models may fail to express the requested emotion reliably, and improving controllability by additional training is costly in both computation and emotion-labeled speech training…

arxiv · Computer Vision

DMA$^2$: Pixel-space Distribution Matching with Adversarial and Anchor Losses

Distribution matching distillation (DMD) provides a general framework for few-step diffusion generation, but its modern text-to-image instantiations have been developed primarily around latent diffusion. It therefore ov…

arxiv · Computer Vision

Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies

Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: differ…

arxiv · Computer Vision

LongLive-Plug: Once-for-All Distillation for Video Generation

Video diffusion models are increasingly developed into specialized models for diverse downstream tasks, and this development often includes a distillation stage, for example to accelerate sampling or to improve long-vid…

arxiv · Computer Vision

PowerSim: Differentiable Physics Simulation and Rendering with Power Diagrams

We introduce PowerSim, a method to bring physically grounded, differentiable dynamics to PowerFoam's power diagram based 3D representation. PowerSim directly couples a pre-trained PowerFoam scene to the Material Point M…

arxiv · Computer Vision

FracGen: Learning How Objects Stretch and Tear with Physics-Informed Video Generation

We introduce FracGen, a fracture-aware video generation model that produces plausible, controllable fracture dynamics from a single image of an intact object, conditioned on physics signals. To train FracGen, we build F…

arxiv · NLP / LLM

Pretraining Latent Information Feedback Transformers with Teacher Supervision

Transformer language models (LMs) are feed-forward: deep-layer representations are never fed back to shallower layers, and the only pathway for information to flow downward across generation steps is the decoded token.…

arxiv · AI

Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning

As agents take on longer and more complex problems, controlling the execution becomes a task in its own right. Each step in the run brings new control choices, like which partial work to build on, whether to start fresh…

arxiv · Computer Vision

LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation

We introduce LIFT, a unified image-to-video generation framework that complements camera control with Layout-In-FuTure control, enabling users to specify what should appear in a future view and where it should appear. T…

arxiv · AI

Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI

Agent performance depends on both reasoning ability and the environment in which it acts. We study test-time AI-for-AI, asking how a Builder can learn to construct better execution environments for a Target while both m…

arxiv · AI

AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation

A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advi…

arxiv · Computer Vision

Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE

Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models. However, conventional token-wise MoE routes tokens independently within a homogeneous expert…

arxiv · NLP / LLM

LongHarness Bench: Stress-Testing Language Model Harnesses for Long-Context Reasoning

Language-model (LM) harnesses enable LMs to operate effectively over long contexts using additional compute. However, existing long-context evaluations are insufficient for distinguishing modern harnesses, reflected by…

arxiv · Computer Vision

CLeaR: A Unified Framework for Resolving the Leakage-Degradation Dilemma in Style Transfer

Style transfer aims to render target content in the style of a reference image, but existing methods often suffer from content leakage, where objects, layouts, or semantics from the style reference appear in the generat…

arxiv · Machine Learning

Multi-Agent Flow Matching with Decoupled Generative Guidance

Generative modeling is widely used for producing diverse objects from complex, multimodal distributions. However, its expressivity does not, in general, come with formal guarantees that the generated objects satisfy har…

arxiv · Machine Learning

Achieving an $O(1/N)$ Optimality Gap in Average-Reward Weakly-Coupled MDPs

We study average-reward weakly-coupled Markov decision processes (WCMDPs), where a WCMDP consists of $N$ smaller MDPs, called arms, that share multiple per-step budget constraints. We consider the setting where the arms…

arxiv · Computer Vision

HelixWorld: A Real-time Interactive Audio-Visual World Model

World simulation is inherently multisensory, demanding synchronized visual and acoustic dynamics in real time. Yet prevailing interactive world models remain strictly silent, focusing exclusively on visual rendering and…

arxiv · Machine Learning

WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms

KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck, we introduce WUSH-KV for low-bit KV-cache quantization. It adapts W…

arxiv · AI

Stochastic World Models for Verifying Vision-Based Neural Feedback Systems

Verifying a vision-based neural feedback system requires a model of the observations its controller acts upon. Such a model must capture the variation the sensor produces, while remaining tractable for closed-loop analy…

hn · Research

Uncensored and Offensive Security AI Models Benchmark

arxiv · Computer Vision

FurE: Efficient Instance-Specific 3D Fur Reconstruction without Animal-Fur Datasets

Realistic and editable animal fur reconstruction from multi-view images is challenging due to fine-scale detail, self-occlusion and obfuscation, and, unlike human hair, the lack of animal-fur datasets. Fur usually cover…

arxiv · NLP / LLM

Telescopic Language Models

One deployed language model must often serve many compute budgets, yet serving each budget still means a separate training or compression run per point. We train a Telescopic Language Model (TLM) to be that continuum: a…

arxiv · Computer Vision

PDMD: Projected Distribution Matching Distillation for Video Diffusion Models

Modern video diffusion models require tens of denoising evaluations over long spatiotemporal token sequences. Distribution Matching Distillation (DMD) reduces the number of function evaluations (NFE) to just a few. Howe…

arxiv · Computer Vision

Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning

Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revisi…

arxiv · NLP / LLM

Retrieving Biblical Intertextual References in Karen Blixen's Seven Gothic Tales

Identifying intertextual references is central to literary scholarship, but computationally difficult when source material is transformed through paraphrase, allusion, historical language, and translation. We investigat…

arxiv · Computer Vision

Reliability-Gated Fusion of Consumer Head and Foot IMUs for Lower-Body 3D Pose

Sparse inertial pose estimation promises camera-free motion capture from consumer devices, but consumer sensors are unreliable: firmware-fused orientations are biased, mounting varies between sessions, and streams drift…

arxiv · Machine Learning

Unifying Distributional Training for One-Step Visual Generation

\emph{Distributional training} provides collective supervision for one-step visual generation by matching real and generated features in frozen representation spaces. We introduce \emph{a unified theoretical framework}…

arxiv · NLP / LLM

Scaling Long-Form Story Generation via Narrative State Tracking

LLMs have demonstrated strong capabilities in creative writing. However, scaling them to full-length novels remains challenging, as maintaining narrative consistency becomes increasingly difficult. Existing story-genera…

arxiv · Machine Learning

TokenCast: Forecasting Token Consumption During LLM Agent Execution

When a large language model (LLM) agent executes the same task, token consumption can vary by over an order of magnitude across runs. The agent chooses its next steps based on tool feedback and intermediate results, whi…

arxiv · Machine Learning

Statistical Learning of Contractive Dynamical Representations for Composite Adaptive Control

We present a representation-learning framework for composite adaptive tracking control under dynamically coupled disturbances. The framework connects classical disturbance-accommodating control (DAC) to recent last-laye…

arxiv · Machine Learning

Neural Harmonic Measure Operator

We introduce Neural Harmonic Measure Operator (NHMO), a neural solver for elliptic PDE problems on variable-shape domains. The harmonic measure of a domain is the boundary probability distribution that, integrated again…

arxiv · Machine Learning

How to Loop MoE: Flatten the Experts, Untie the Attention

Looped Transformers reuse one block of layers several times: by spending extra computation they push a model of fixed size further, and so use its parameters more fully; while sparse mixture-of-experts (MoE) models acti…

arxiv · Machine Learning

KV-streams for Efficient Compaction in Agentic Reinforcement Learning

Scaling the horizon of agentic LLMs is bottlenecked by the need to fit ever longer context traces in GPU memory. Context compaction has been the most popular mechanism to alleviate this issue, keeping GPU memory constan…

arxiv · NLP / LLM

Towards Communication-Efficient Social Intelligence in Language Agents

Socially intelligent language agents must negotiate, coordinate, and resolve conflicting preferences while respecting the time and attention of both participants. Balancing these demands is challenging because agents mu…

arxiv · NLP / LLM

Improving Test-Time Scaling with Adaptive Looped Transformers

Looped transformers have demonstrated promising parameter efficiency by reusing layers for latent computation. Prior studies compare looped and non-looped models at matched parameters or per-token FLOPs. However, to the…

arxiv · Computer Vision

Copy the Same, Distill the Difference: Initializing Linear Vision Transformers

Linear Vision Transformers (ViTs) are designed to replace the attention in Softmax ViTs with the linear-complexity attention operator for more efficient token routing, but they require from-scratch pre-training and typi…

arxiv · AI

FinAutoRubric: Expert-Guided Automatic Rubric Generation for Evaluating Financial Research Agents

Evaluating finance research agents requires rubrics that reflect expert standards and fix the values correct as of an information cutoff. Expert-reviewed finance benchmarks rely on fixed, per-item rubrics, which are cos…

arxiv · Computer Vision

InfiniHand: Streaming World-Space Hand Motion Estimation from Egocentric Video

World-space hand motion estimation from egocentric video requires recovering 3D articulated hand geometry while tracking camera egomotion. Existing approaches heavily rely on cascading independent hand pose estimators a…

arxiv · AI

Shockingly Simple Self-retrospection Improves Agentic Models Without RL

People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future action…

arxiv · NLP / LLM

Harness Learning Enables Generalizable Test-Time Adaptation

A language-model agent is jointly defined by its model and its harness, the executable program that organizes model calls, tool use, and information flow. Because different tasks call for different ways of organizing th…

arxiv · Computer Vision

GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space

Novel view synthesis from sparse images must reconcile faithful reconstruction of observed regions with plausible completion of unseen content, while maintaining world consistency across viewpoints. Existing geometry-ba…

arxiv · AI

Failure-Transparent Agents: Benchmarking Post-Failure Reporting in Tool-Using Language Models

Tool-using agents can fail twice: a required tool can fail, and the agent can then report success without the evidence needed to justify it. Existing benchmarks often entangle this reporting failure with tool selection,…

arxiv · Computer Vision

FlowAct-R2: Beyond Talking Avatar via Streaming Multimodal References and Proactive Agent Planning

We present FlowAct-R2, a framework for interactive humanoid video generation that combines continuous multimodal control with proactive agent planning. Our method consists of two coupled components. First, a Streaming M…

arxiv · Computer Vision

Impact of Patient Orientation in Single- and Multi-View Camera Environments for AI-based Rehabilitation Monitoring

Automated quality assessment of rehabilitation exercises relies heavily on accurate human pose estimation from video data. Although numerous RGB-based pose estimation methods have been proposed, the impact of camera pla…

arxiv · Computer Vision

Superquadric Primitive Decomposition of 3D point clouds via Geometric-Aware Inlier Refinement

The decomposition of 3D point clouds into interpretable geometric primitives remains a longstanding challenge in Computer Vision and Computer Graphics. Among the available representations, superquadrics offer a compact…

arxiv · Computer Vision

Hard Vision, Easy Vision: What GPT-6 Astra Reveals Across Computer Vision

Frontier general-purpose systems are rapidly expanding beyond visual understanding into capabilities traditionally handled by dedicated computer-vision models. As these capabilities expand, a central question for the co…

arxiv · Machine Learning

X-Reset: Scaling Object-Centric Reinforcement Learning via Cross-Embodiment Resets

Reinforcement learning (RL) in simulation can train dexterous manipulation policies without robot demonstrations, but training a single generalist policy with task-agnostic rewards faces a severe exploration problem: ap…

arxiv · Computer Vision

Lagrangian--Hamiltonian Flows for Video Prediction and Image Generation: A Symplectic Perspective

We introduce LHFM, a geometric framework for learning image dynamics. Drawing on structures central to classical mechanics, symplectic geometry, and geometric quantization, LHFM represents each image as an exact Lagrang…

arxiv · Computer Vision

Mind the RefGAP: Correcting Reference Attention in Diffusion-Based Visual Editing

Reference-guided diffusion editors struggle to faithfully reproduce user-provided references. We identify a potential bottleneck in diffusion editors: many methods provide limited reference-attention allocation. For exa…

Browse briefing issues →

Get the briefing

The one story that matters, 5 headlines and the paper everyone's citing — every Tuesday, free.

Subscribe free