All Articles
17650 articles total
Pre-Flight: A Benchmark for Evaluating Large Language Models on Aviation Operational Knowledge
Introduction of Pre-Flight, an open-source benchmark for evaluating LLM reasoning and safety in aviation operational knowledge.
Actual causality in fault trees
A theoretical study applying Halpern & Pearl's theory of actual causality to fault trees to improve failure diagnostics in complex systems.
CLAP: Closed-Loop Training, Evaluation, and Release Control for Domain Agent Post-training
Proposal of CLAP, a closed-loop method for post-training domain agents to improve data quality and release stability in manufacturing scenarios.
Safety Targeted Embedding Exploit via Refinement
Research demonstrating a gradient-guided attack (STEER) that bypasses LLM safety filters by translating refusal-triggering words into low-resource languages.
CamoNAS: Neural Architecture Search for Enhanced Camouflaged Object Detection
Introduction of CamoNAS, a frequency-aware Neural Architecture Search framework for enhancing the detection of camouflaged objects.
SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use
Introduction of SkillCoach, a framework that uses self-evolving rubrics to evaluate and improve how LLM agents utilize operational skills.
Spec-AUF: Accept-Until-Fail Training under Train-Inference Misalignment for Masked Block Drafters
Introduction of Spec-AUF, a training objective for masked block drafters in speculative decoding to improve token emission length and generation speed.
Rethinking Complexity Metrics for LLM-Integrated Applications: Beyond Source Code
Presentation of HECATE, a tool that assesses complexity in both the prompt and code layers of LLM-integrated applications.
ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair
Introduction of ContextSniper, a token-efficient code memory layer that reduces context usage and costs for repository-level program repair agents.
Underwater Suit-Wearing Cyborg Insect Capable of Diving and Terra-Aqua Travel
Researchers have developed a cyborg insect capable of diving and navigating both terrestrial and aquatic environments.
The Safari MCP server for web developers
A new Model Context Protocol (MCP) server has been released specifically for Safari, aimed at helping web developers.
Path-level Hindsight Instructions for Semantic Exploration in Vision-Language Navigation
Phi-Nav introduces a framework for Vision-Language Navigation (VLN) that uses hindsight reasoning to align instructions with an agent's actual exploration.
Mastermind: Strategy-grounded Learning for Repository-Scale Vulnerability Reproduction
Mastermind is a dual-loop framework that separates transferable strategy learning from task-specific experience to improve repository-scale vulnerability reproduction.
SimWorlds: A Multi-Agent System for Dynamic 3D Scene Creation
SimWorlds is a multi-agent framework that generates dynamic, editable 4D scenes from text using Blender-specific procedural knowledge.
Repair the Amplifier, Not the Symptom: Stable World-Model Correction for Agent Rollouts
WM-SAR proposes a world-model corrector that identifies causal subgraphs to repair failed planning graphs in place, improving agent rollouts.
Verifiable Knowledge Expansion through Retrieval-Grounded Formal Concept Analysis
A new framework uses Formal Concept Analysis (FCA) and a retrieval-grounded SLM oracle for verifiable knowledge expansion in ontology construction.
Subliminal Clocks: Latent Time Modelling in Diffusion Language Models
Research shows that Diffusion Language Models internally encode a latent representation of the diffusion timestep, which can be used to modulate model confidence.
Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification
Vera is an automated safety testing framework for LLM agents that uses a three-stage pipeline to discover risks and verify them through evidence-grounded predicates.
MMIR-TCM: Memory-Integrated Multimodal Inference and Retrieval for TCM Clinical Decision Support
MMIR-TCM integrates MLLMs with memory-augmented segmentation and RAG to provide clinical decision support for Traditional Chinese Medicine diagnosis.
Every AI Visibility Tool Is Lying to You
A discussion regarding the potential inaccuracies and deceptive nature of existing AI visibility and observability tools.