All Articles
17800 articles total
Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models?
Researchers introduce Drop-Then-Recovery (DTR) to analyze architectural redundancy in Vision-Language-Action (VLA) models, finding that language backbones are often highly redundant for robotic tasks.
RS-Diffuser: Risk-Sensitive Diffusion Planning with Distributional Value Guidance
RS-Diffuser is a risk-sensitive offline diffusion planning framework that uses distributional value critics to allow flexible risk-averse or risk-seeking behaviors in robotic navigation.
Improving Adversarial Robustness via Activation Amplification and Attenuation
The A3 module improves adversarial robustness in neural networks by jointly learning to amplify and attenuate signals through a lightweight activation scaling mechanism.
Output-Space Allocation Costs for Calibration-Guided LLM Compression: An Empirical Study
An empirical study on LLM compression explores whether aligning allocation costs with output-space objectives improves fidelity, noting a tradeoff between accuracy and perplexity.
SHIFT: Gate-Modulated Activation Steering for Knowledge Conflict Mitigation in Retrieval-Augmented Generation
SHIFT is a framework that uses gate-modulated activation steering to resolve knowledge conflicts between retrieved context and parametric knowledge in RAG systems.
NLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window Adaptation
A training-free method for selecting full-attention layers in hybrid models using Negative Log-Likelihood (NLL) guidance to optimize long-context inference efficiency.
Position Bias Correction is Insufficient for One-Pass Attention Sorting
Research indicates that simple position bias correction is insufficient to match the performance of iterative attention sorting for long-context language models.
Optimizing Teacher-Student Partitioning for Scalable Knowledge Distillation on HPC Systems
A new HPC-aware methodology for knowledge distillation decouples teacher and student partitioning to achieve up to 67% higher throughput than the TRL library.
Parameter-Efficient Quantum-Inspired Fast Weight Programmers for Traffic-Matrix Forecasting
The paper proposes QKAN-FWPs, quantum-inspired recurrent models that provide efficient traffic-matrix forecasting with significantly lower memory and training budgets.
Pepti-drift: Toxicity-Repulsive Drifting for Antigen-Conditioned Discrete Peptide Generation
Pepti-drift is a toxicity-aware latent refinement framework designed to generate antigen-specific binding peptides while avoiding toxicity.
Replacing Systemd with OpenRC in Debian
A discussion on the process and implications of replacing systemd with OpenRC in Debian Linux distributions.
1.38 Millimeter Microcontroller
Presentation of a microcontroller with an extremely small physical footprint of 1.38 millimeters.
US Grid Constraints: Towards 40GW+ of Behind-the-Meter Datacenter by 2028?
An analysis of US power grid constraints and the potential rise of behind-the-meter datacenter capacity by 2028.
Do Speech Emphasis Models Generalize across Languages and Emotions?
Introduction of the MMEE corpus, a multilingual multi-emotion emphasis dataset for improving speech emphasis detection models.
Enhancing Numerical Prediction in LLMs via Smooth MMD Alignment
Proposes Smooth Maximum Mean Discrepancy (SMMD) to improve the numerical precision of LLM predictions using value-distance kernels.
Bifocal Diffusion Language Models: Asymmetric Bidirectional Context for Parallel Generation
Introduces Bifocal Diffusion Language Models and R2LM, combining causal attention with Mamba SSMs to improve parallel generation throughput.
KG2Cypher: Data-Centric Pipeline for Building Enterprise Text-to-Cypher Systems
Presents KG2Cypher, a data-centric pipeline for converting natural language questions into executable Cypher queries for enterprise knowledge graphs.
End-to-End Dynamic Sparsity for Resource-Adaptive LLM Inference
Proposes the L2A framework for resource-adaptive LLM inference, allowing models to dynamically adjust computation based on runtime budgets.
Flexformer: Flexible Linear Transformer with Learnable Attention Kernel
Introduces Flexformer, a linear Transformer that utilizes learnable attention kernels to reduce complexity while maintaining expressiveness.
From General-Purpose Audio Tagging to Spatially Grounded Sound Event Localization and Detection
Explores the AT2SELD framework for extending general-purpose audio tagging models to spatially grounded sound event localization and detection.