IntentScore: Intent-Conditioned Action Evaluation for Computer-Use Agents

When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don’t

arXiv:2604.06422v1 Announce Type: cross Abstract: Understanding when Vision-Language Models (VLMs) will behave unexpectedly, whether models can reliably predict their own behavior, and if models adhere

Bi-Level Optimization for Single Domain Generalization

arXiv:2604.06349v1 Announce Type: cross Abstract: Generalizing from a single labeled source domain to unseen target domains, without access to any target data during training, remains

Uncertainty Estimation for Deep Reconstruction in Actuatic Disaster Scenarios with Autonomous Vehicles

arXiv:2604.06387v1 Announce Type: cross Abstract: Accurate reconstruction of environmental scalar fields from sparse onboard observations is essential for autonomous vehicles engaged in aquatic monitoring. Beyond

$S^3$: Stratified Scaling Search for Test-Time in Diffusion Language Models

arXiv:2604.06260v1 Announce Type: cross Abstract: Test-time scaling investigates whether a fixed diffusion language model (DLM) can generate better outputs when given more inference compute, without

Harnessing Hyperbolic Geometry for Harmful Prompt Detection and Sanitization

arXiv:2604.06285v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have become essential for tasks such as image synthesis, captioning, and retrieval by aligning textual and visual