All Articles
16503 articles total
Opto-ViT-v2: Noise-Resilient On-Chip Fine-Tuning for Photonic Near-Sensor Vision Transformer Accelerators
Opto-ViT-v2 is a framework for parameter-efficient fine-tuning on photonic near-sensor Vision Transformer accelerators, significantly reducing storage and energy costs while resisting noise.
JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language Models
JailMeter is an evidence-based evaluation framework for measuring the effectiveness of jailbreak attacks on LLMs, including a distilled small language model version for efficiency.
Making Single-Cell Data Distillation Auditable: Traceable Real-Cell Coresets via Discrete Min-Max Selection
Research on traceable single-cell data distillation using Minmax-CF, which allows reducing dataset size while maintaining the ability to trace results back to original measured cells.
ChannelGuard: Safe Models Do Not Compose into Safe Multi-Agent Systems
ChannelGuard is a training-free defense framework that implements information-bottleneck gates on inter-agent channels in multi-agent LLM systems to block prompt injection and tool poisoning.
BRIM: Workload-Balanced Dual-Sided Bit-Serial Sparse Inference Accelerator
BRIM is a hardware-software co-designed sparse inference accelerator that uses Cyclic-Balanced Pruning and Pairwise Slot Donation to solve workload imbalance in bit-serial accelerators.
ChainWatch: A Kill Chain-Aligned Sequential Detection Framework for Multi-Step Attacks in MCP-Based AI Agent Systems
ChainWatch is a sequential detection framework that uses Hidden Markov Models and a kill-chain approach to identify multi-step attacks in AI agent systems using the Model Context Protocol (MCP).
Stateful Guardrails for Multi-Turn LLM Systems: A Conversational Risk Accumulation Framework
The authors introduce a framework to detect Conversational Risk Accumulation (CRA) in LLMs, where harmless turns in a conversation build up into a harmful outcome.
Economic Evaluations of Language Models
EconEvals is an open-source evaluation suite designed to measure the economic value and labor-market impact of LLMs across various US occupations.
Challenges of Explainability in Continual Learning for Time Series Forecasting
This research explores the use of explainability techniques like Grad-CAM and attention rollout to understand continual learning in time series forecasting.
Scale-Aware Learning of Chaotic Dynamics on Unstructured Meshes via Binned Spectral Losses
The study presents a method for scale-aware learning of chaotic dynamics on unstructured meshes using binned spectral losses and graph-Laplacian frequency bands.
Simulating Eutopia: Revisiting Long-term Fairness with Outcomes, Performativity, and Dynamics
The paper introduces Eutopia, a lending-process simulator used to study and improve long-term fairness and equity in AI-driven decision makers.
LAARA: Layer-Aware Adaptive Rank Allocation for Parameter-Efficient Fine-Tuning
LAARA is a search-free framework for parameter-efficient fine-tuning that dynamically allocates ranks to transformer layers using Fisher estimates.
Decodable but Not Detectable: A Leakage Fingerprint for Near-OOD Benchmarks
The authors identify a 'leakage fingerprint' to detect when OOD benchmarks are contaminated with training data, which can lead to misleadingly high performance scores.
Cross-Subject Semantic Decoding with Shared-Space Alignment for Generalized Neural Representation Learning
A new framework for cross-subject semantic decoding aligns neural responses to speech perception into a shared latent space to improve generalization across different individuals.
From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation
Prefix-GRPO is a reinforcement learning framework that improves small-model agents by decomposing teacher trajectories into replayed prefixes and online continuations.
Leveraging Offline Supervision for Efficient and Generalizable Reinforcement Learning in Large-Scale Vision-Language-Action Models
This work investigates hybrid offline-online training for Vision-Language-Action (VLA) models, showing that offline supervision can double training efficiency while maintaining OOD performance.
Escape IntelliJ: Scala and Kotlin LSPs on Emacs Eglot
Discussion on configuring Scala and Kotlin Language Server Protocol (LSP) support in Emacs using the Eglot package.
Cruller: Bun's Zig Runtime, Continued on Zig 0.16
Updates on Cruller, a Zig-based runtime for Bun, now continuing development on Zig version 0.16.
Global Difference Constraint Propagation for Constraint Programming
A research paper proposing a global propagator for difference constraints in constraint programming to improve solving speed and completeness.
Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model
Research on steering representations of materials science mechanisms within the open-weight Gemma-4 model to understand how LLMs represent physical laws.