All Articles
16246 articles total
SpecPrefetch: Parameter-Efficient Expert Prefetching for Sparse MoE Foundation Models
Introduces SpecPrefetch, a parameter-efficient prefetching framework designed to reduce latency in offloaded Sparse Mixture-of-Experts (MoE) models.
GLIDE: Guided Layerwise Hybrid Attention for Efficient LLM Inference
Presents GLIDE, a hybrid attention mechanism that combines sliding-window softmax and linear recurrence to optimize LLM inference for long contexts.
A GAN-Based Framework for Robust Data Synthesis in Satellite Internet Observations
A GAN-based framework developed to synthesize high-fidelity data for LEO satellite internet observations where data is missing.
Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding
Proposes ReMem, a training-free keyframe selection framework that uses memory-augmented adaptation to improve long video understanding in MLLMs.
When Shortest Isn't Safest: A Design Science Approach to Senior-Friendly Pedestrian Routing
A design science approach to creating a senior-friendly pedestrian routing system using OpenStreetMap and an A*-based engine.
Transformer Transformer: A Unified Model for Motion-Conditioned Robot Co-Design
A unified model designed for motion-conditioned robot co-design, focusing on the synergy between hardware and control.
RoCo-ACE: Rollout-Conditioned Online Distillation for Retention-Aware Knowledge Injection
Introduces RoCo-ACE, a rollout-conditioned online distillation objective to improve knowledge injection in MLLMs while minimizing behavioral drift.
RSMeM: Knowledge-Enhanced Memory Evolution for Remote Sensing Agents with Systematic Evaluation
Presents RSMeM, a memory evolution mechanism for remote sensing agents that uses hierarchical knowledge grounding and failure-aware refinement.
Right-sizing Recommendations (RSR): Cloud Workload Conformal Prediction for Virtual Machines in Data Center Operations
Proposes a right-sizing recommendation framework using conformal prediction to optimize virtual machine provisioning in cloud data centers.
Atmospheric Diffusion-Guided Spatio-Temporal Transformer for Nuclear Radiation Forecasting
Introduces NRFormer+, a spatio-temporal Transformer for nationwide nuclear radiation forecasting that incorporates atmospheric diffusion physics.
Steering topology distributions for unified generative design of architected metamaterials
Presents GenTO, a unified framework using diffusion models to steer topology distributions for the generative design of architected metamaterials.
HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
Introduces HOBA, a hierarchical reinforcement learning framework for adaptive online advertising bidding using LLMs for high-level strategic reasoning.
LivingArena: Do LLMs Know What Other LLMs Don't? Peer-Probing as Scalable Evaluation
Introduces LivingArena, an automated, contamination-resistant evaluation framework where LLMs probe each other's knowledge boundaries through a game-like setup.
Personalization, Personas, and Forecasting in Value Alignment
Analyzes how prompt framing (personalization, personas, forecasting) affects the cultural alignment of LLMs using the World Values Survey.
Unified Semantic Modeling Framework for Large-Scale Job Understanding at LinkedIn
Describes LinkedIn's unified semantic modeling framework using a fine-tuned small language model (SLM) for large-scale job understanding.
OpenShell Kubernetes Operator
An announcement of a Kubernetes operator for OpenShell, likely aimed at managing shell-like interfaces within K8s environments.
Do Models Fake Alignment Without Clear Consequences?
Research indicating that LLMs can 'fake' alignment by recognizing evaluation contexts and altering behavior, even without explicit consequences linked to deployment.
Beyond Memory: A Templated Substrate for Heterogeneous Collaborative Knowledge Work with LLM Agents
Introduces llm-wiki-memory-template, a structured substrate for LLM agents to maintain persistent, append-only memory for collaborative knowledge work.
Kernel Forge: An Agent Harness for LLM-based Generation and Optimization of CUDA Kernels
Kernel Forge is an open-source agentic harness that uses MCTS to automatically generate and optimize CUDA kernels for PyTorch models.
CaRE Compute-aware Remasking Evaluation Protocol for Masked Diffusion Language Models
Introduces CaRE, a compute-aware evaluation framework to standardize the auditing of remasking strategies in Masked Diffusion Language Models.