All Articles
16503 articles total
Fine-grained Computation-Communication Overlap via Tile-level Signaling and Scheduling for Mixture-of-Experts
Introduces a fine-grained computation-communication overlap technique for Mixture-of-Experts (MoE) models, achieving up to 2.64x end-to-end speedup on multi-GPU systems.
Juxtaposition of Shallow Reservoir-Triggered Seismicity and Deep Tectonic Locking in the Qiaojia-Dongchuan Seismic Gap
Analyzes seismic gaps using a decoupling model to distinguish between shallow reservoir-triggered seismicity and deep tectonic strain.
Causal dictionary learning reveals and validates transcription-factor binding features in genomic language models
Develops a causal dictionary learning framework to extract and validate interpretable transcription-factor binding features in genomic language models.
SCPP: A Unified Python Library for Soft Clustering
Presents SCPP, an open-source Python library for soft clustering that provides a standardized scikit-learn-compatible interface for various algorithms.
Understanding Developer Pain Points in Federated Learning: Insights from Stack Overflow and GitHub
An empirical study of developer pain points in Federated Learning based on analysis of Stack Overflow and GitHub, highlighting gaps in tooling and documentation.
Adaptive Capitulation: A Structural Failure Mode of LLM Responses in Vulnerability Contexts
Identifies 'adaptive capitulation' as a failure mode in LLMs where the model validates social injustice before facilitating discouraged behavior in sensitive contexts.
Protecting our FLOSS commons from LLMs
A discussion on protecting Free and Libre Open Source Software (FLOSS) from being exploited by LLMs.
Building Trust in Autonomous Commerce: A Verifiable Global Event Timeline and AI-Ready Fraud Intelligence Layer
Proposes a verifiable global event timeline and fraud intelligence layer for autonomous agentic commerce to ensure auditability and temporal ordering.
BaseRT: Advancing Best-in-Class LLM Inference with Apple M5 Neural Accelerators
Introduces BaseRT, a native Metal inference runtime for LLMs on Apple M5 chips that significantly outperforms llama.cpp and MLX in throughput.
Unlearning as Distribution Restoration: A Controlled Counterfactual Study, a Validated Selective Screen, and the Limits of Oracle-Free Certification
A study on machine unlearning, proposing a base-anchored held-out screen to better audit the removal of specific knowledge from models.
Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents
Audits tool-augmented LLM agents, finding that safety-oriented system prompts can paradoxically increase 'unfaithful safety refusals' when tools fail silently.
REGEN: Replay-recycling for Expert-to-Generalist distillation with Offline Reinforcement Learning
Introduces REGEN, a method using offline RL and replay-recycling from teacher models to distill generalist LLMs at a lower computational cost.
Predictive Extrema, Unprofitable Policies: An AI-Assisted Audit of Candle-Based Binance Spot Timing Models
An audit of AI-assisted candle-based trading models on Binance Spot, concluding they are generally unprofitable after costs.
MoA-Structured Decode Attention DNF Derivation, KV-Cache Accumulation, GQA/MQA, and OpenACC Kernel
Derives memory-optimal inference artifacts for transformer attention, including a C/OpenACC GPU kernel and KV-cache optimizations.
ModPack: An Extensible Teleoperation Interface for Bimanual Mobile Manipulation
Presents ModPack, an extensible teleoperation interface for bimanual mobile manipulation robots, with open-sourced hardware and software.
Integrity of peer-to-peer distributed LLM inference under malicious nodes
Proposes a method to detect malicious nodes in P2P distributed LLM inference by using secret canary inputs to measure activation drift.
ANSI escape injection in MCP servers: Hidden from humans, visible to AI
Discussion on how ANSI escape sequences can be used for prompt injection in Model Context Protocol (MCP) servers to hide malicious instructions from humans while remaining visible to AI.
Recovering Clinical Utility Under Differential Privacy: Empirical Validation of Adaptive Federated Aggregation on Heterogeneous Cardiovascular Datasets
Introduction of the FedCVR framework for federated learning on real cardiovascular datasets, demonstrating that adaptive optimization can preserve clinical utility under differential privacy.
Structured Latent Space Modeling over Multi-Scale Temporal Patches for Multivariate Time Series Forecasting
Presentation of M2Patch, a CNN-based architecture for multivariate time series forecasting that uses multi-scale patching and latent space constraints for better accuracy and linear complexity.
Auditing Retrieval-Augmented LLM Hypotheses for Longitudinal Cell Painting Morphology
A retrieval-augmented interpretation framework for longitudinal Cell Painting morphology that uses LLMs to generate auditable biological hypotheses with quantitative validation tests.