All Articles
16483 articles total
SafeStep: AI-powered Travel Assistance for Elderly People with Frailty or Dementia
SafeStep is an AI-driven travel assistance system for elderly people with dementia, combining LLMs with behavioral prediction to generate personalized failure scenarios and interventions.
Safeguards for Speech2Speech LLM-Assistants: A Case Study in Automotive Applications
A study evaluates the effectiveness of transcript-based and tool-based safeguards for speech-to-speech LLM assistants in cars, finding both insufficient due to latency and non-deterministic behavior.
Explaining Weather Bulletins via ILP
This research utilizes Inductive Logic Programming (ILP) via the FastLAS2 framework to generate interpretable hypotheses explaining weather bulletins from raw meteorological data.
Differentiable Logic Programming to Mitigate Reasoning Shortcuts in Neurosymbolic Systems
A new method using matrix-based differentiable logic programming is proposed to reduce reasoning shortcuts in neurosymbolic systems, ensuring more robust concept mapping and constraint satisfaction.
Identifying Good Rules for Efficient SAT Encodings of Single-Constant Multiplication Using Machine Learning
A neuro-symbolic framework using Graph Neural Networks is introduced to accelerate SAT encoding for the Single Constant Multiplication problem in hardware design.
Bound-Founded Semantics for Answer Set Programming with Difference Constraints: Preliminary Report
This report introduces a many-sorted variant of the Bound-founded Logic of Here-and-There (HTb) to provide a unified semantic foundation for Answer Set Programming with linear constraints.
Delivery, Not Storage: Cue-Anchored Working Memory as a Harness Property for Coding Agents
The paper proposes a cue-anchored working memory model for coding agents to ensure reliability by treating memory as a harness property rather than an agent choice.
Beyond Independent Optimization: Compression, MoE Routing, and Quantization Interactions in Multimodal Edge Intelligence
This research examines the complex interactions between compression, MoE routing, and quantization in multimodal edge intelligence, arguing they should not be treated as independent optimizations.
GuardianAgentBench: Where Agents Fail and How to Guard Them
The authors introduce GuardianAgentBench, a benchmark to evaluate the safety and reliability of LLM agents across various domains and adversarial attack modes.
Workflow-Localized Mechanism Learning: Attribution-Guided Repair and Knowledge Reuse for Structured Agent Skills
The paper introduces Workflow-Localized Mechanism Learning (WML) to help agents repair and reuse knowledge for structured procedural skills.
Naju: A Native Discrete State-Space Model with Independent Retention and Writing for Long-Sequence Memory
Naju is a native discrete state-space model designed for long-sequence memory tracking by decoupling retention and writing mechanisms.
Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers
This study investigates the trustworthiness and stability of LLM-generated summaries by proposing a two-level diagnostic protocol.
EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization
EmoAgent-R1 uses reinforcement learning and dynamic agent specialization to improve multimodal emotion understanding in MLLMs.
HiMe: Real-Time Self-Hosted Personal Agent Platform for Health Insights with Wearable Devices
HiMe is a privacy-first, locally deployable platform for real-time health monitoring using wearable device data and LLM agents.
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs
The researchers present Faster IndexTTS-2, an acceleration method for autoregressive text-to-speech models using NVIDIA TensorRT and TensorRT-LLM.
Can Generative Recommendation Reach Cold Items? A Temporal Perspective on Semantic-ID Generation
This paper analyzes the limitations of current semantic-ID based generative recommendation systems in handling cold-start items.
Claude Cookbook
Anthropic provides a 'Cookbook' of examples and guides for developers using Claude.
Code Monitor Red Teaming for Public-Test-Passing Code
Researchers introduce CodeMonitorBench to test if LLM verifiers can identify bugs in code that has already passed public tests.
Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions
The MisKnow-Agent framework reveals that Deep Research LLM agents are vulnerable to misleading knowledge, often adopting false conclusions in final reports.
Source-Prior-Driven Selective Adaptation for Efficient Diffusion Model Finetuning
A new source-prior-driven selective adaptation method for diffusion models improves the trade-off between target-specific generation and the retention of general capabilities.