CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning

arXiv:2512.02551v2 Announce Type: replace-cross Abstract: In this paper, we propose CUDA-L2, a system that combines large language models (LLMs) and reinforcement learning (RL) to automatically

MiniScope: A Least Privilege Framework for Authorizing Tool Calling Agents

arXiv:2512.11147v1 Announce Type: cross Abstract: Tool calling agents are an emerging paradigm in LLM deployment, with major platforms such as ChatGPT, Claude, and Gemini adding

Reducing Fragmentation and Starvation in GPU Clusters through Dynamic Multi-Objective Scheduling

arXiv:2512.10980v1 Announce Type: cross Abstract: GPU clusters have become essential for training and deploying modern AI systems, yet real deployments continue to report average utilization

KathDB: Explainable Multimodal Database Management System with Human-AI Collaboration

arXiv:2512.11067v1 Announce Type: cross Abstract: Traditional DBMSs execute user- or application-provided SQL queries over relational data with strong semantic guarantees and advanced query optimization, but

Agile Flight Emerges from Multi-Agent Competitive Racing

arXiv:2512.11781v1 Announce Type: cross Abstract: Through multi-agent competition and the sparse high-level objective of winning a race, we find that both agile flight (e.g., high-speed

TimeScope: Towards Task-Oriented Temporal Grounding In Long Videos

December 9, 2025

arXiv:2509.26360v3 Announce Type: replace-cross
Abstract: Identifying key temporal intervals within long videos, known as temporal grounding (TG), is important to video understanding and reasoning tasks. In this paper, we introduce a new form of the temporal grounding problem, textbfTask-oriented Temporal Grounding (textbfToTG), which is driven by the requirements of downstream tasks rather than explicit time-interval descriptions. For example, a ToTG input may be “explain why the man in the video is sent to the hospital,” whereas traditional TG would take an explicit temporal description such as “the moments when the man is tripped by a stone and falls to the ground.” This new ToTG formulation presents significant challenges for existing TG methods, as it requires jointly performing deep task comprehension and fine-grained temporal localization within long videos. To address these challenges, we conduct a systematic set of studies. First, we construct textbfa new benchmark ToTG-Bench, which comprehensively evaluates ToTG performance across diverse settings. Second, we introduce textbfa new temporal-ground method TimeScope, which performs coarse-to-fine localization through a progressive reasoning process. Leveraging extensive supervised fine-tuning with carefully curated chain-of-thought (CoT) data from a variety of scenarios, TimeScope generalizes effectively across tasks and domains. Our evaluation demonstrates textbfTimeScope’s empirical advantages over existing baselines from three perspectives: (1) substantial improvements in grounding precision, (2) significant benefits to downstream tasks, and (3) strong generalizability across different scenarios. All models, datasets, and source code will be fully open-sourced to support future research in this area.

Subscribe for Updates

Copyright 2025 dijee Intelligence Ltd. dijee Intelligence Ltd. is a private limited company registered in England and Wales at Media House, Sopers Road, Cuffley, Hertfordshire, EN6 4RY, UK registeration number 16808844