< 返回
研究论文
- 《Beyond Imitation: Learning Safe End-to-End Autonomous Driving from Hard Negatives》2026-07-14Existing imitation learning methods for end-to-end autonomous driving predominantly learn from successful demonstrations by minimizing geometric deviations from expert trajectories. This paradigm implicitly assumes that…
- 《ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval》2026-07-14Leveraging Multimodal Large Language Models (MLLMs) via contrastive learning has become a mainstream paradigm for improving the performance of Universal Multimodal Retrieval (UMR). However, previous works have ignored t…
- 《DriveVA: Video Action Models are Zero-Shot Drivers》2026-07-14Generalization is a central challenge in autonomous driving, as real-world deployment requires robust performance under unseen scenarios, sensor domains, and environmental conditions. Recent world model-based planning m…
- 《TIGER: Taming Identity, Geometry, and Generative Priors for High-Quality Face Video Restoration》2026-07-14Face Video Restoration (FVR) aims to recover high-fidelity facial videos from degraded input while preserving identity and semantic consistency across frames. Existing methods often struggle to simultaneously address th…
- 《GAIA: A Data Flywheel System for Training GUI Test-Time Scaling Critic Models》2026-07-14While Large Vision-Language Models (LVLMs) have significantly advanced GUI agents’ capabilities in parsing textual instructions, interpreting screen content, and executing tasks, a critical challenge persists: the irrev…
- 《UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation》2026-07-14In-Image Machine Translation (IIMT) aims to translate scene text in an image and render the translated text back into the original regions while preserving the overall visual appearance. Recent unified multimodal models…
- 《CausalDrive: Real-time Causal World Models for Autonomous Driving》2026-07-14World models have emerged as a promising paradigm for scaling autonomous driving (AD) data, yet existing video generative models fall short as interactive simulators. Layout-conditioned renderers rely on "oracle" future…
- 《Pondering the Way: Spatial-perceiving World Action Model for Embodied Navigation》2026-07-14Existing world model-based planners for visual navigation typically follow a verification-centric paradigm, decoupling goal intent from trajectory synthesis. This approach suffers from candidate dependence, heavy comput…
- 《Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation》2026-07-14Reinforcement learning has emerged as a principled post-training paradigm for Temporal Video Grounding (TVG) due to its on-policy optimization, yet existing GRPO-based methods remain fundamentally constrained by sparse …
- 《Structured Progressive Knowledge Activation for LLM-Driven Neural Architecture Search》2026-07-14This paper focuses on a key challenge in Neural Architecture Search (NAS): integrating established architectural knowledge while exploring new designs under expensive evaluations. Large language models (LLMs) are a prom…
- 《Beyond Absolute Scores: Relative Edit-induced Difference for Generalizable Image Aesthetic Assessment》2026-07-14Traditional Image Aesthetic Assessment (IAA) methods mainly rely on regressing absolute Mean Opinion Scores (MOS). However, such a paradigm overlooks the inherently dynamic nature of human aesthetic perception, which re…
- 《Restoring Exploration after Post-Training: Latent Exploration Decoding for Large Reasoning Models》2026-07-14Large Reasoning Models (LRMs) have recently achieved strong mathematical and code reasoning performance through Reinforcement Learning (RL) post-training. However, we show that modern reasoning post-training induces an …
- 《Time Series Reasoning via Process-Verifiable Thinking Data Synthesis and Scheduling for Tailored LLM Reasoning》2026-07-14Time series is a pervasive data type across various application domains, rendering the reasonable solving of diverse time series tasks a long-standing goal. Recent advances in large language models (LLMs), especially th…
- 《MECAT:AMulti-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks》2026-07-14While large audio-language models have advanced open-ended audio understanding, they still fall short of nuanced human-level comprehension. This gap persists largely because current benchmarks, limited by data annotatio…
- 《OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models》2026-07-14We present OmniVoice, a massive multilingual zero-shot text-to-speech (TTS) model that scales to over 600 languages. At its core is a novel diffusion language model-style discrete non-autoregressive (NAR) architecture. …
- 《Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously》2026-07-14Online Video Large Language Models (VideoLLMs) play a critical role in supporting responsive, real-time interaction. Existing methods focus on streaming perception, lacking a synchronized logical reasoning stream. Howev…
- 《CoME: Empowering Channel-of-Mobile-Experts with Informative Hybrid-Capabilities Reasoning》2026-07-14Mobile Agents can autonomously execute user instructions, which requires hybrid-capabilities reasoning, including screen summary, subtask planning, action decision and action function. However, existing agents struggle …
- 《DashengTokenizer: One layer is enough for unified audio understanding and generation》2026-06-24This paper introduces DashengTokenizer, a continuous audio tokenizer engineered for joint use in both understanding and generation tasks. Unlike conventional approaches, which train acoustic tokenizers and subsequently …
- 《The Interspeech 2026 Audio Encoder Capability Challenge for Large Audio Language Models》2026-06-24This paper presents the Interspeech 2026 Audio Encoder Capability Challenge, a benchmark specifically designed to evaluate and advance the performance of pre-trained audio encoders as front-end modules for Large Audio L…
- 《MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning》2026-07-14Current Vision-Language-Action (VLA) paradigms in autonomous driving primarily rely on Imitation Learning (IL), which introduces inherent challenges such as distribution shift and causal confusion. Online Reinforcement …
- 《DriveFine: Refining-Augmented Masked Diffusion VLA for Precise and Robust Driving》2026-07-14Vision-Language-Action (VLA) models for autonomous driving increasingly adopt generative planners trained with imitation learning followed by reinforcement learning. Diffusion-based planners suffer from modality alignme…
- 《Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension》2026-07-14Existing LLM test-time scaling laws emphasize the emergence of self-reflective behaviors through extended reasoning length. Nevertheless, this vertical scaling strategy often encounters plateaus in exploration as the mo…
- 《SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization》2026-06-24Multimodal large language models (MLLMs) have demonstrated impressive reasoning and instruction-following capabilities, yet their expanded modality space introduces new compositional safety risks that emerge from comple…
- 《FutureMind: Equipping Small Language Models with Strategic Thinking-Pattern Priors via Adaptive Knowledge Distillation》2026-06-24Small Language Models (SLMs) are attractive for cost-sensitive and resourcelimited settings due to their efficient, low-latency inference. However, they often struggle with complex, knowledge-intensive tasks that requir…
- 《Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers》2026-07-14Reinforcement learning (RL) has emerged as a crucial approach for enhancing the capabilities of large language models. However, in Mixture-of-Experts (MoE) models, the routing mechanism often introduces instability, eve…
- 《GLCLAP: A Novel Contrastive Learning Pre-trained Model for Contextual Biasing in ASR》2026-06-24Automatic Speech Recognition (ASR) that supports prompts has shown remarkable versatility. For contextual biasing with these systems, a pivotal factor lies in obtaining well-matched prompts. To address this issue, Contr…
- 《Controllable Pedestrian Video Editing for Multi-View Driving Scenarios via Motion Sequence》2026-06-24Pedestrian detection models in autonomous driving systems often lack robustness due to insufficient representation of dangerous pedestrian scenarios in training datasets. To address this limitation, we present a novel f…
- 《Zipvoice: Fast and high-quality zero-shot text-to-speech with flow matching》2026-07-14Existing large-scale zero-shot text-to-speech (TTS) models deliver high speech quality but suffer from slow inference speeds due to massive parameters. To address this issue, this paper introduces ZipVoice, a high-quali…
- 《Not All Pixels Are Equal: Learning Pixel Hardness for Semantic Segmentation》2026-06-24Semantic segmentation has witnessed great progress. Despite the impressive overall results, the segmentation performance in some hard areas (e.g., small objects or thin parts) is still not promising. A straightforward s…
- 《Portrait Shadow Removal Using Context-Aware Illumination Restoration Network》2026-06-24Portrait shadow removal is a challenging task due to the complex surface of the face. Although existing work in this field makes substantial progress, these methods tend to overlook information in the background areas. …
- 《Cr-ctc: Consistency regularization on ctc for improved speech recognition》2026-06-30Connectionist Temporal Classification (CTC) is a widely used method for automatic speech recognition (ASR), renowned for its simplicity and computational efficiency. However, it often falls short in recognition performa…
- 《Zipformer: A faster and better encoder for automatic speech recognition》2026-06-30The Conformer has become the most popular encoder model for automatic speech recognition (ASR). It adds convolution modules to a transformer to learn both local and global dependencies. In this work we describe a faster…
- 《Pruned RNN-T for fast, memory-efficient ASR training》2026-06-30The RNN-Transducer (RNN-T) framework for speech recognition has been growing in popularity, particularly for deployed real-time ASR systems, because it combines high accuracy with naturally streaming recognition. One of…