AI/ML arXiv cs.AI

EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization

Introduces EvoThink, a framework that uses self-pruning and preference optimization to reduce redundant verification steps in Large Reasoning Models (LRMs).

AI/ML arXiv cs.AI

The Giant Hippocampus: From Structural Monoculture to a System of Systems

Argues that the AI field's reliance on Transformers is a structural error and proposes Heterogeneous Topological Networks to better mirror the biological cortex.

AI/ML arXiv cs.AI

Coordinating from Memory: Graph-Structured Experience Reuse for Multi-Agent Adaptation in Dynamic Manufacturing

Proposes the Graph-Structured Experiential Memory (GSEM) framework to improve multi-agent coordination in dynamic manufacturing through relational graphs.

AI/ML arXiv cs.AI

CLARK: Closed-loop Learning for Adaptive Reasoning over Knowledge Graphs

Introduces CLARK, a framework integrating knowledge graphs and symbolic rule mining for adaptive reasoning and classification under uncertainty.

Software Engineering arXiv cs.AI

Safe Remediation as Risk-Constrained Intervention Decision in Microservice Systems

Reformulates automated microservice remediation as a risk-constrained intervention problem using Constrained Markov Decision Processes (CMDP).

Hardware/Chips arXiv cs.AI

EvoDRC: A Self-Evolving Agentic Framework for Automated DRC Violation Repair

Presents EvoDRC, an agentic framework that evolves repair skills to automate Design Rule Check (DRC) violations in advanced-node physical design.

Cybersecurity Hacker News

Local AI that finds sensitive files on your Mac before attackers do

An AI-powered tool designed for macOS to help users identify and secure sensitive files before they can be exploited by attackers.

Software Engineering Hacker News

Why malloc always does more than I asked for?

A technical exploration of how malloc allocates more memory than requested due to alignment and padding requirements.

Tech Business/VC TechCrunch

ServiceNow bets $40 million on Indian banking software specialist to expand its financial services push

ServiceNow invests $40 million in Indian banking software firm BusinessNext to expand its AI-driven financial services capabilities.

Other Ars Technica

Orcas team up to ram sunfish until they explode

Observations of orcas ramming sunfish as a predatory or playful behavior.

Hardware/Chips arXiv cs.AI

Symbol and Footprint Database for Electronic Components by Agentic Recognition and Generation

SFgen is an agentic recognition and generation system that creates symbols and footprints for electronic components to automate PCB design.

AI/ML arXiv cs.AI

Silent Failures in Multimodal Agentic Search:A Diagnostic Taxonomy and Cross-Judge Evaluation

Researchers identify 'silent failures' in multimodal agentic search and propose a diagnostic taxonomy and evaluation pipeline to measure true trajectory correctness.

AI/ML arXiv cs.AI

Rewarding Better Thinking for LLM Preference Alignment

Introduces Thinking Checklist Reward (TCR), a process-oriented reward system for RL-based preference alignment to improve LLM reasoning trajectories.

Cybersecurity arXiv cs.AI

Know Your Agent: Reconnaissance-Driven Pentesting of AI Agents

Presents Know Your Agent (KYA), a framework for reconnaissance-driven pentesting of AI agents to identify and mitigate indirect prompt injection attacks.

AI/ML arXiv cs.AI

DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations

DocOps is a verifiable benchmark for evaluating the ability of autonomous agents to perform complex document operations and maintain global consistency.

AI/ML arXiv cs.AI

JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety

Introduces Janus, a foresight-oriented safety framework that trains guard models to anticipate and block delayed risks in long-horizon agent trajectories.

AI/ML arXiv cs.AI

Mitigating Scaffolding Collapse in Socratic Tutors via Representation Alignment

Researchers propose a framework to prevent Socratic tutors from abandoning guided inquiry and revealing answers directly under student pressure by using representation alignment.

AI/ML arXiv cs.AI

Euclean: Automated Geometry Problem Formalization with Unified Verification in Lean

The Euclean framework automates geometry problem formalization in Lean using a four-stage process, creating the largest geometry formalization dataset in Lean.

Cybersecurity arXiv cs.AI

CrackedPDFs: A Controlled Benchmark for Hidden Prompt Injection in PDFs

The CrackedPDFs benchmark introduces a controlled method to evaluate hidden prompt injection vulnerabilities in PDF documents used by LLM systems.

AI/ML arXiv cs.AI

HyGRL: Adaptive Hybrid Graph Reasoning for Multi-Entity Questions

HyGRL is a unified framework that integrates unstructured text with structured knowledge graphs via imitation and reinforcement learning for better multi-entity reasoning.