All Articles
17790 articles total
NUMA: Cores, memory, and the distance between them
A discussion on NUMA (Non-Uniform Memory Access) architecture, focusing on the relationship between CPU cores and memory distance.
Hippocampus-DETR: An Explicit Memory Object Detection Framework Based on Hippocampus Modeling
Introduction of Hippocampus-DETR, an object detection framework that simulates biological hippocampal memory for better generalization and data efficiency.
WattLayer: Get Layers Right to Estimate Inference Energy of Neural Networks
WattLayer proposes a layer-wise energy estimation model to accurately measure the inference energy consumption of neural networks across different hardware.
Applicability of memorization indicators for early spotting of overfitting while recalibrating sEMG-decoders on low sample sizes
Research on using ReLU activation statistics as memorization indicators to spot overfitting early during sEMG-decoder recalibration with low sample sizes.
GNBAN: Graph Neural Basis Attention Networks for Long-Horizon Forecasting over Large Entity Sets
GNBAN is a Graph Neural Basis Attention Network designed for scalable and interpretable long-horizon forecasting of large entity sets in retail.
S$^2$-VLA: State-Space Guided Vision-Language-Action Models for Long-Horizon Manipulation
S^2-VLA introduces a state-space guided adaptive attention mechanism to improve long-horizon robotic manipulation by dynamically fusing visual and language features.
SpatialUAV: Benchmarking Spatial Intelligence for Low-Altitude UAV Perception, Collaboration, and Motion
SpatialUAV provides a benchmark for evaluating spatial intelligence in low-altitude UAVs, focusing on 3D inference and multi-view collaboration.
A Study of Temporal Fusion Strategies for Named Entity Recognition in Historical Texts
A study on using temporal metadata and fusion strategies to improve Named Entity Recognition (NER) in historical texts.
SEADA: An efficient methodology for optimizing mixed-precision DNNs on multi-precision spatial architectures
SEADA is a methodology for optimizing mixed-precision DNNs on multi-precision spatial architectures to reduce latency and energy footprint.
Triadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMs
Triadic Werewolf evaluates LLM Theory-of-Mind by introducing a 'Jester' role to test multi-agent reasoning across conflicting utility functions.
Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models?
Researchers introduce Drop-Then-Recovery (DTR) to analyze architectural redundancy in Vision-Language-Action (VLA) models, finding that language backbones are often highly redundant for robotic tasks.
RS-Diffuser: Risk-Sensitive Diffusion Planning with Distributional Value Guidance
RS-Diffuser is a risk-sensitive offline diffusion planning framework that uses distributional value critics to allow flexible risk-averse or risk-seeking behaviors in robotic navigation.
Improving Adversarial Robustness via Activation Amplification and Attenuation
The A3 module improves adversarial robustness in neural networks by jointly learning to amplify and attenuate signals through a lightweight activation scaling mechanism.
Output-Space Allocation Costs for Calibration-Guided LLM Compression: An Empirical Study
An empirical study on LLM compression explores whether aligning allocation costs with output-space objectives improves fidelity, noting a tradeoff between accuracy and perplexity.
SHIFT: Gate-Modulated Activation Steering for Knowledge Conflict Mitigation in Retrieval-Augmented Generation
SHIFT is a framework that uses gate-modulated activation steering to resolve knowledge conflicts between retrieved context and parametric knowledge in RAG systems.
NLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window Adaptation
A training-free method for selecting full-attention layers in hybrid models using Negative Log-Likelihood (NLL) guidance to optimize long-context inference efficiency.
Position Bias Correction is Insufficient for One-Pass Attention Sorting
Research indicates that simple position bias correction is insufficient to match the performance of iterative attention sorting for long-context language models.
Optimizing Teacher-Student Partitioning for Scalable Knowledge Distillation on HPC Systems
A new HPC-aware methodology for knowledge distillation decouples teacher and student partitioning to achieve up to 67% higher throughput than the TRL library.
Parameter-Efficient Quantum-Inspired Fast Weight Programmers for Traffic-Matrix Forecasting
The paper proposes QKAN-FWPs, quantum-inspired recurrent models that provide efficient traffic-matrix forecasting with significantly lower memory and training budgets.
Pepti-drift: Toxicity-Repulsive Drifting for Antigen-Conditioned Discrete Peptide Generation
Pepti-drift is a toxicity-aware latent refinement framework designed to generate antigen-specific binding peptides while avoiding toxicity.