AI/ML arXiv cs.AI

Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows

Analysis of LLM integration in fraud detection and trust-and-safety workflows, identifying a critical lack of operational evidence regarding latency and cost.

AI/ML arXiv cs.AI

Inference Economics of Enterprise Coding Agents: A Case Study of Cloud vs. On-Premise LLMs

A case study comparing API-based vs. on-premise quantized LLMs for coding agents, revealing that while on-premise can save TCO, it increases developer debugging burden.

Cybersecurity arXiv cs.AI

SingGuard-NSFA: Extensible Guardrails for Agentic AI via Generative Reasoning and Real-Time Classification

Introduction of SingGuard-NSFA (nsfaguard), a guardrail framework using generative reasoning and classification to secure agentic AI against operational threats.

Cybersecurity arXiv cs.AI

Baselines Before Architecture: Evaluating Coding Agents for Autonomous Penetration Testing

A study evaluating autonomous penetration testing agents, arguing that plain coding agents often perform as well as complex architectures when matched by model version.

AI/ML arXiv cs.AI

Self-Improving AI Coding Agents Through Accumulated Behavioral Rules: A Closed-Loop Framework

A closed-loop framework for AI coding agents that converts human review feedback into persistent behavioral rules to prevent recurring mistakes without weight updates.

AI/ML arXiv cs.AI

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models

A privacy-aware edge-cloud collaborative inference framework for LLMs that uses authenticated KV caches to reduce latency and payload size.

AI/ML arXiv cs.AI

Final Authority in AI Governance: Frontier-Provider Sovereignty and Action-Centered Deployer Governance

A research paper comparing governance models for AI systems, arguing that final authority over high-impact actions should reside with the deployer rather than the model provider.

AI/ML arXiv cs.AI

LessonBench-V1: A Benchmark Dataset for Evaluating AI Lesson Generation Agents

Introduction of LessonBench-V1, a benchmark dataset for evaluating AI agents that generate educational STEM content based on pedagogical methodologies.

AI/ML arXiv cs.AI

Beyond Backbone Backpropagation: A Decoupled Strategy for Efficient Transfer Learning

Proposes a decoupled training strategy for efficient transfer learning in CNNs and Transformers that reduces computational cost and CO2 emissions.

AI/ML arXiv cs.AI

The Perplexity Trap: When Patent Law Makes Human Writing Look Like AI

Analyzes the failure of open-source AI detectors to distinguish between LLM-generated and human-written patent texts due to structural linguistic similarities.

AI/ML arXiv cs.AI

Federated Explainable Artificial Intelligence: Roles, Architectures, Evaluation, and Open Challenges

A comprehensive survey of Federated Explainable AI (FedXAI), exploring the integration of privacy-preserving learning and model transparency.

AI/ML arXiv cs.AI

Uncertainty-Aware Sequential Decision Rules for Event-Triggered LLM Invocation in Streaming Systems

Presents a risk-based framework for deciding when to invoke expensive LLMs in streaming systems to balance semantic understanding with computational cost.

AI/ML arXiv cs.AI

Autonomous UAV Route Planning for Coverage Maximization in Environmental Monitoring: A Systematic Literature Review

A systematic literature review on autonomous UAV route planning for maximizing environmental monitoring coverage under various constraints.

AI/ML arXiv cs.AI

Compaction as Epistemic Failure: How Agentic LLM Tools Fabricate Confirmed Results from Killed Processes

Documents a critical failure in Claude Code where timed-out process outputs are erroneously recorded as confirmed results in session summaries.

AI/ML arXiv cs.AI

HRO: Hierarchical Room-to-Object Framework for Zero-Shot Object Goal Navigation with Large Language Models

Introduces the HRO framework for zero-shot object goal navigation, using LLMs to provide hierarchical spatial reasoning for better robot exploration.

Other arXiv cs.AI

When is the combined load identifiable from a stress-intensity profile? A coupled forward-inverse study on SIFBench finite-element data

A study on the identifiability of combined loads from stress-intensity profiles using the SIFBench finite-element dataset.

Software Engineering Hacker News

Dense Arena Interning: The Engine of Compiler Performance

A technical discussion on how dense arena interning can drive compiler performance improvements.

Tech Business/VC The Verge

OnePlus officially gives up on the US and Europe

OnePlus has officially announced it will cease product launches in the US and European markets.

AI/ML arXiv cs.AI

Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0

Researchers investigate whether agent-optimization methods can compound performance over time in continual learning scenarios using Terminal-Bench 2.0.

AI/ML arXiv cs.AI

AI-accelerated End-to-End Framework for Rapid Professional Upskilling

A new framework uses AI to accelerate various stages of professional upskilling, from content development to assessment.