Explainable AI in kidney stone detection and segmentation: a mini review

Kidney stones are one of the most common renal disorders that can produce severe complications if not diagnosed and treated early. Recently, advances in AI

Feasibility testing of a home-based exercise intervention in children with cerebral palsy who are ambulant—a study protocol of the HOME-EX study

Children gain increased health and well-being by participating in physical activity. Children with cerebral palsy who are ambulatory (CP-A) are known to be less physically

Patient and clinician perceptions, expectations, and usability of ankle exoskeletons for daily living: a mixed-methods survey study

Ankle exoskeletons offer promising support for individuals with chronic foot drop, yet user and clinician perspectives on their use in daily living remain underexplored. Related

Why digital health fails silently: a sociotechnical theory of health information technology–related risk

IntroductionHealth information technology (HIT) is now integral to healthcare delivery, supporting clinical documentation, prescribing, diagnostics, and care coordination. Although these technologies offer substantial benefits, they

A maturity model framework for federated networks of trusted research environments

IntroductionA Trusted Research Environment (TRE) is a highly secure computer system where sensitive data is stored that researchers can access remotely and make use of

When to Stop Reusing: Dynamic Gradient Gating for Sample-Efficient RLVR

May 20, 2026

arXiv:2605.19425v1 Announce Type: cross
Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become the dominant paradigm for advanced reasoning in Large Language Models (LLMs), but rollout samples are expensive to obtain, making sample efficiency a critical bottleneck. A natural remedy is to reuse each rollout batch for multiple gradient updates, a standard practice in classical RL. Yet in RLVR, this amplifies policy shift, leading to severe performance degradation. Detecting the onset of degradation early enough to stop reuse remains an open and challenging problem. We close this gap by identifying the textitDisproportionate Weight Divergence (DWD) phenomenon: performance degradation is synchronized with a sharp surge in the textttlm_head weight change, while intermediate layers remain stable. Empirically, we verify that DWD emerges consistently across diverse LLMs and tasks. Theoretically, we prove that (i) harmful gradients concentrate at the textttlm_head while intermediate layers are structurally attenuated, and (ii) the textttlm_head gradient norm lower-bounds the policy divergence. These results establish the textttlm_head gradient norm as a principled, real-time signal of catastrophic policy shift. Guided by this insight, we propose textitDynamic Gradient Gating (DGG), a lightweight intervention that monitors the textttlm_head gradient norm in real time and intercepts harmful gradients before they corrupt the optimizer. DGG consistently matches or exceeds the standard single-use baseline, achieving up to $2.93times$ sample efficiency and $2.14times$ wall-clock speedup across math, ALFWorld, WebShop, and search-augmented QA tasks.

Subscribe for Updates

Copyright 2025 dijee Intelligence Ltd. dijee Intelligence Ltd. is a private limited company registered in England and Wales at Media House, Sopers Road, Cuffley, Hertfordshire, EN6 4RY, UK registration number 16808844