All Articles
16950 articles total
Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows
Analysis of LLM integration in fraud detection and trust-and-safety workflows, identifying a critical lack of operational evidence regarding latency and cost.
Inference Economics of Enterprise Coding Agents: A Case Study of Cloud vs. On-Premise LLMs
A case study comparing API-based vs. on-premise quantized LLMs for coding agents, revealing that while on-premise can save TCO, it increases developer debugging burden.
SingGuard-NSFA: Extensible Guardrails for Agentic AI via Generative Reasoning and Real-Time Classification
Introduction of SingGuard-NSFA (nsfaguard), a guardrail framework using generative reasoning and classification to secure agentic AI against operational threats.
Baselines Before Architecture: Evaluating Coding Agents for Autonomous Penetration Testing
A study evaluating autonomous penetration testing agents, arguing that plain coding agents often perform as well as complex architectures when matched by model version.
Self-Improving AI Coding Agents Through Accumulated Behavioral Rules: A Closed-Loop Framework
A closed-loop framework for AI coding agents that converts human review feedback into persistent behavioral rules to prevent recurring mistakes without weight updates.
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models
A privacy-aware edge-cloud collaborative inference framework for LLMs that uses authenticated KV caches to reduce latency and payload size.
Final Authority in AI Governance: Frontier-Provider Sovereignty and Action-Centered Deployer Governance
A research paper comparing governance models for AI systems, arguing that final authority over high-impact actions should reside with the deployer rather than the model provider.
LessonBench-V1: A Benchmark Dataset for Evaluating AI Lesson Generation Agents
Introduction of LessonBench-V1, a benchmark dataset for evaluating AI agents that generate educational STEM content based on pedagogical methodologies.
Beyond Backbone Backpropagation: A Decoupled Strategy for Efficient Transfer Learning
Proposes a decoupled training strategy for efficient transfer learning in CNNs and Transformers that reduces computational cost and CO2 emissions.
The Perplexity Trap: When Patent Law Makes Human Writing Look Like AI
Analyzes the failure of open-source AI detectors to distinguish between LLM-generated and human-written patent texts due to structural linguistic similarities.
Federated Explainable Artificial Intelligence: Roles, Architectures, Evaluation, and Open Challenges
A comprehensive survey of Federated Explainable AI (FedXAI), exploring the integration of privacy-preserving learning and model transparency.
Uncertainty-Aware Sequential Decision Rules for Event-Triggered LLM Invocation in Streaming Systems
Presents a risk-based framework for deciding when to invoke expensive LLMs in streaming systems to balance semantic understanding with computational cost.
Autonomous UAV Route Planning for Coverage Maximization in Environmental Monitoring: A Systematic Literature Review
A systematic literature review on autonomous UAV route planning for maximizing environmental monitoring coverage under various constraints.
Compaction as Epistemic Failure: How Agentic LLM Tools Fabricate Confirmed Results from Killed Processes
Documents a critical failure in Claude Code where timed-out process outputs are erroneously recorded as confirmed results in session summaries.
HRO: Hierarchical Room-to-Object Framework for Zero-Shot Object Goal Navigation with Large Language Models
Introduces the HRO framework for zero-shot object goal navigation, using LLMs to provide hierarchical spatial reasoning for better robot exploration.
When is the combined load identifiable from a stress-intensity profile? A coupled forward-inverse study on SIFBench finite-element data
A study on the identifiability of combined loads from stress-intensity profiles using the SIFBench finite-element dataset.
Dense Arena Interning: The Engine of Compiler Performance
A technical discussion on how dense arena interning can drive compiler performance improvements.
OnePlus officially gives up on the US and Europe
OnePlus has officially announced it will cease product launches in the US and European markets.
Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0
Researchers investigate whether agent-optimization methods can compound performance over time in continual learning scenarios using Terminal-Bench 2.0.
AI-accelerated End-to-End Framework for Rapid Professional Upskilling
A new framework uses AI to accelerate various stages of professional upskilling, from content development to assessment.