Hardware/Chips arXiv cs.AI

Opto-ViT-v2: Noise-Resilient On-Chip Fine-Tuning for Photonic Near-Sensor Vision Transformer Accelerators

Opto-ViT-v2 is a framework for parameter-efficient fine-tuning on photonic near-sensor Vision Transformer accelerators, significantly reducing storage and energy costs while resisting noise.

Cybersecurity arXiv cs.AI

JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language Models

JailMeter is an evidence-based evaluation framework for measuring the effectiveness of jailbreak attacks on LLMs, including a distilled small language model version for efficiency.

AI/ML arXiv cs.AI

Making Single-Cell Data Distillation Auditable: Traceable Real-Cell Coresets via Discrete Min-Max Selection

Research on traceable single-cell data distillation using Minmax-CF, which allows reducing dataset size while maintaining the ability to trace results back to original measured cells.

Cybersecurity arXiv cs.AI

ChannelGuard: Safe Models Do Not Compose into Safe Multi-Agent Systems

ChannelGuard is a training-free defense framework that implements information-bottleneck gates on inter-agent channels in multi-agent LLM systems to block prompt injection and tool poisoning.

Hardware/Chips arXiv cs.AI

BRIM: Workload-Balanced Dual-Sided Bit-Serial Sparse Inference Accelerator

BRIM is a hardware-software co-designed sparse inference accelerator that uses Cyclic-Balanced Pruning and Pairwise Slot Donation to solve workload imbalance in bit-serial accelerators.

Cybersecurity arXiv cs.AI

ChainWatch: A Kill Chain-Aligned Sequential Detection Framework for Multi-Step Attacks in MCP-Based AI Agent Systems

ChainWatch is a sequential detection framework that uses Hidden Markov Models and a kill-chain approach to identify multi-step attacks in AI agent systems using the Model Context Protocol (MCP).

AI/ML arXiv cs.AI

Stateful Guardrails for Multi-Turn LLM Systems: A Conversational Risk Accumulation Framework

The authors introduce a framework to detect Conversational Risk Accumulation (CRA) in LLMs, where harmless turns in a conversation build up into a harmful outcome.

AI/ML arXiv cs.AI

Economic Evaluations of Language Models

EconEvals is an open-source evaluation suite designed to measure the economic value and labor-market impact of LLMs across various US occupations.

AI/ML arXiv cs.AI

Challenges of Explainability in Continual Learning for Time Series Forecasting

This research explores the use of explainability techniques like Grad-CAM and attention rollout to understand continual learning in time series forecasting.

AI/ML arXiv cs.AI

Scale-Aware Learning of Chaotic Dynamics on Unstructured Meshes via Binned Spectral Losses

The study presents a method for scale-aware learning of chaotic dynamics on unstructured meshes using binned spectral losses and graph-Laplacian frequency bands.

AI/ML arXiv cs.AI

Simulating Eutopia: Revisiting Long-term Fairness with Outcomes, Performativity, and Dynamics

The paper introduces Eutopia, a lending-process simulator used to study and improve long-term fairness and equity in AI-driven decision makers.

AI/ML arXiv cs.AI

LAARA: Layer-Aware Adaptive Rank Allocation for Parameter-Efficient Fine-Tuning

LAARA is a search-free framework for parameter-efficient fine-tuning that dynamically allocates ranks to transformer layers using Fisher estimates.

AI/ML arXiv cs.AI

Decodable but Not Detectable: A Leakage Fingerprint for Near-OOD Benchmarks

The authors identify a 'leakage fingerprint' to detect when OOD benchmarks are contaminated with training data, which can lead to misleadingly high performance scores.

AI/ML arXiv cs.AI

Cross-Subject Semantic Decoding with Shared-Space Alignment for Generalized Neural Representation Learning

A new framework for cross-subject semantic decoding aligns neural responses to speech perception into a shared latent space to improve generalization across different individuals.

AI/ML arXiv cs.AI

From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation

Prefix-GRPO is a reinforcement learning framework that improves small-model agents by decomposing teacher trajectories into replayed prefixes and online continuations.

AI/ML arXiv cs.AI

Leveraging Offline Supervision for Efficient and Generalizable Reinforcement Learning in Large-Scale Vision-Language-Action Models

This work investigates hybrid offline-online training for Vision-Language-Action (VLA) models, showing that offline supervision can double training efficiency while maintaining OOD performance.

Software Engineering Hacker News

Escape IntelliJ: Scala and Kotlin LSPs on Emacs Eglot

Discussion on configuring Scala and Kotlin Language Server Protocol (LSP) support in Emacs using the Eglot package.

Software Engineering Hacker News

Cruller: Bun's Zig Runtime, Continued on Zig 0.16

Updates on Cruller, a Zig-based runtime for Bun, now continuing development on Zig version 0.16.

AI/ML arXiv cs.AI

Global Difference Constraint Propagation for Constraint Programming

A research paper proposing a global propagator for difference constraints in constraint programming to improve solving speed and completeness.

AI/ML arXiv cs.AI

Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model

Research on steering representations of materials science mechanisms within the open-weight Gemma-4 model to understand how LLMs represent physical laws.