AI/ML arXiv cs.AI

Pre-Flight: A Benchmark for Evaluating Large Language Models on Aviation Operational Knowledge

Introduction of Pre-Flight, an open-source benchmark for evaluating LLM reasoning and safety in aviation operational knowledge.

AI/ML arXiv cs.AI

Actual causality in fault trees

A theoretical study applying Halpern & Pearl's theory of actual causality to fault trees to improve failure diagnostics in complex systems.

AI/ML arXiv cs.AI

CLAP: Closed-Loop Training, Evaluation, and Release Control for Domain Agent Post-training

Proposal of CLAP, a closed-loop method for post-training domain agents to improve data quality and release stability in manufacturing scenarios.

Cybersecurity arXiv cs.AI

Safety Targeted Embedding Exploit via Refinement

Research demonstrating a gradient-guided attack (STEER) that bypasses LLM safety filters by translating refusal-triggering words into low-resource languages.

AI/ML arXiv cs.AI

CamoNAS: Neural Architecture Search for Enhanced Camouflaged Object Detection

Introduction of CamoNAS, a frequency-aware Neural Architecture Search framework for enhancing the detection of camouflaged objects.

AI/ML arXiv cs.AI

SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use

Introduction of SkillCoach, a framework that uses self-evolving rubrics to evaluate and improve how LLM agents utilize operational skills.

AI/ML arXiv cs.AI

Spec-AUF: Accept-Until-Fail Training under Train-Inference Misalignment for Masked Block Drafters

Introduction of Spec-AUF, a training objective for masked block drafters in speculative decoding to improve token emission length and generation speed.

Software Engineering arXiv cs.AI

Rethinking Complexity Metrics for LLM-Integrated Applications: Beyond Source Code

Presentation of HECATE, a tool that assesses complexity in both the prompt and code layers of LLM-integrated applications.

AI/ML arXiv cs.AI

ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair

Introduction of ContextSniper, a token-efficient code memory layer that reduces context usage and costs for repository-level program repair agents.

Hardware/Chips Hacker News

Underwater Suit-Wearing Cyborg Insect Capable of Diving and Terra-Aqua Travel

Researchers have developed a cyborg insect capable of diving and navigating both terrestrial and aquatic environments.

Software Engineering Hacker News

The Safari MCP server for web developers

A new Model Context Protocol (MCP) server has been released specifically for Safari, aimed at helping web developers.

AI/ML arXiv cs.AI

Path-level Hindsight Instructions for Semantic Exploration in Vision-Language Navigation

Phi-Nav introduces a framework for Vision-Language Navigation (VLN) that uses hindsight reasoning to align instructions with an agent's actual exploration.

Cybersecurity arXiv cs.AI

Mastermind: Strategy-grounded Learning for Repository-Scale Vulnerability Reproduction

Mastermind is a dual-loop framework that separates transferable strategy learning from task-specific experience to improve repository-scale vulnerability reproduction.

AI/ML arXiv cs.AI

SimWorlds: A Multi-Agent System for Dynamic 3D Scene Creation

SimWorlds is a multi-agent framework that generates dynamic, editable 4D scenes from text using Blender-specific procedural knowledge.

AI/ML arXiv cs.AI

Repair the Amplifier, Not the Symptom: Stable World-Model Correction for Agent Rollouts

WM-SAR proposes a world-model corrector that identifies causal subgraphs to repair failed planning graphs in place, improving agent rollouts.

AI/ML arXiv cs.AI

Verifiable Knowledge Expansion through Retrieval-Grounded Formal Concept Analysis

A new framework uses Formal Concept Analysis (FCA) and a retrieval-grounded SLM oracle for verifiable knowledge expansion in ontology construction.

AI/ML arXiv cs.AI

Subliminal Clocks: Latent Time Modelling in Diffusion Language Models

Research shows that Diffusion Language Models internally encode a latent representation of the diffusion timestep, which can be used to modulate model confidence.

Cybersecurity arXiv cs.AI

Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification

Vera is an automated safety testing framework for LLM agents that uses a three-stage pipeline to discover risks and verify them through evidence-grounded predicates.

AI/ML arXiv cs.AI

MMIR-TCM: Memory-Integrated Multimodal Inference and Retrieval for TCM Clinical Decision Support

MMIR-TCM integrates MLLMs with memory-augmented segmentation and RAG to provide clinical decision support for Traditional Chinese Medicine diagnosis.

AI/ML Hacker News

Every AI Visibility Tool Is Lying to You

A discussion regarding the potential inaccuracies and deceptive nature of existing AI visibility and observability tools.