All Articles
16513 articles total
EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization
Introduces EvoThink, a framework that uses self-pruning and preference optimization to reduce redundant verification steps in Large Reasoning Models (LRMs).
The Giant Hippocampus: From Structural Monoculture to a System of Systems
Argues that the AI field's reliance on Transformers is a structural error and proposes Heterogeneous Topological Networks to better mirror the biological cortex.
Coordinating from Memory: Graph-Structured Experience Reuse for Multi-Agent Adaptation in Dynamic Manufacturing
Proposes the Graph-Structured Experiential Memory (GSEM) framework to improve multi-agent coordination in dynamic manufacturing through relational graphs.
CLARK: Closed-loop Learning for Adaptive Reasoning over Knowledge Graphs
Introduces CLARK, a framework integrating knowledge graphs and symbolic rule mining for adaptive reasoning and classification under uncertainty.
Safe Remediation as Risk-Constrained Intervention Decision in Microservice Systems
Reformulates automated microservice remediation as a risk-constrained intervention problem using Constrained Markov Decision Processes (CMDP).
EvoDRC: A Self-Evolving Agentic Framework for Automated DRC Violation Repair
Presents EvoDRC, an agentic framework that evolves repair skills to automate Design Rule Check (DRC) violations in advanced-node physical design.
Local AI that finds sensitive files on your Mac before attackers do
An AI-powered tool designed for macOS to help users identify and secure sensitive files before they can be exploited by attackers.
Why malloc always does more than I asked for?
A technical exploration of how malloc allocates more memory than requested due to alignment and padding requirements.
ServiceNow bets $40 million on Indian banking software specialist to expand its financial services push
ServiceNow invests $40 million in Indian banking software firm BusinessNext to expand its AI-driven financial services capabilities.
Orcas team up to ram sunfish until they explode
Observations of orcas ramming sunfish as a predatory or playful behavior.
Symbol and Footprint Database for Electronic Components by Agentic Recognition and Generation
SFgen is an agentic recognition and generation system that creates symbols and footprints for electronic components to automate PCB design.
Silent Failures in Multimodal Agentic Search:A Diagnostic Taxonomy and Cross-Judge Evaluation
Researchers identify 'silent failures' in multimodal agentic search and propose a diagnostic taxonomy and evaluation pipeline to measure true trajectory correctness.
Rewarding Better Thinking for LLM Preference Alignment
Introduces Thinking Checklist Reward (TCR), a process-oriented reward system for RL-based preference alignment to improve LLM reasoning trajectories.
Know Your Agent: Reconnaissance-Driven Pentesting of AI Agents
Presents Know Your Agent (KYA), a framework for reconnaissance-driven pentesting of AI agents to identify and mitigate indirect prompt injection attacks.
DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations
DocOps is a verifiable benchmark for evaluating the ability of autonomous agents to perform complex document operations and maintain global consistency.
JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety
Introduces Janus, a foresight-oriented safety framework that trains guard models to anticipate and block delayed risks in long-horizon agent trajectories.
Mitigating Scaffolding Collapse in Socratic Tutors via Representation Alignment
Researchers propose a framework to prevent Socratic tutors from abandoning guided inquiry and revealing answers directly under student pressure by using representation alignment.
Euclean: Automated Geometry Problem Formalization with Unified Verification in Lean
The Euclean framework automates geometry problem formalization in Lean using a four-stage process, creating the largest geometry formalization dataset in Lean.
CrackedPDFs: A Controlled Benchmark for Hidden Prompt Injection in PDFs
The CrackedPDFs benchmark introduces a controlled method to evaluate hidden prompt injection vulnerabilities in PDF documents used by LLM systems.
HyGRL: Adaptive Hybrid Graph Reasoning for Multi-Entity Questions
HyGRL is a unified framework that integrates unstructured text with structured knowledge graphs via imitation and reinforcement learning for better multi-entity reasoning.