Other arXiv cs.AI

Programming Language Policy as an AI Literacy Equity Problem: A 15-Nation Comparative Analysis

A 15-nation analysis examines how national CS education policies and language choices create structural inequities in AI literacy.

AI/ML arXiv cs.AI

PRISM Edit: One Vector for All Temporal Answers

PRISM Edit is a model editing framework that handles temporal facts by optimizing polysemous representations without architectural changes.

Software Engineering arXiv cs.AI

Fail-Aware and Explainable Test Oracle Prediction

FOCAL is a code LLM-based discriminative oracle predictor that identifies whether test prefixes will pass or fail to improve fault detection.

Software Engineering Hacker News

An unusual way for your DHCP server to run out of dynamic IPs

A discussion regarding an unusual scenario where a DHCP server exhausts its pool of dynamic IP addresses.

AI/ML arXiv cs.AI

BeatEdit: Symbolic Music Generation as Explicit Editing

BeatEdit introduces a framework for symbolic music generation that uses explicit edit operations on a beat-grid-anchored representation rather than generating from scratch.

Cybersecurity arXiv cs.AI

AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation

AMT-X is a phase-structured multi-turn red-teaming framework designed to better evaluate LLM safety by using a multi-role jury and gated checklists.

AI/ML arXiv cs.AI

RepTran: Search-Based Repair of Transformer Models

RepTran is a search-based repair method for Transformer models that targets feed-forward networks to identify and optimize suspicious weights.

AI/ML arXiv cs.AI

ProgramTab: Boosting Table Reasoning of LLMs via Programmatic Paradigm

ProgramTab improves LLM table reasoning by using a programmatic paradigm that employs Python code for preprocessing and extraction.

AI/ML arXiv cs.AI

HandFlow: Fully Generative 4D Hand Recovery with Flow Matching

HandFlow is a generative flow-matching framework for 4D hand recovery from monocular video, achieving high temporal smoothness and speed.

AI/ML arXiv cs.AI

DeepBias: Adaptive In-depth Probing of Social Biases in LVLMs

DeepBias is an adaptive framework for probing social biases in Large Vision-Language Models using a dynamic generation-evolution-probing loop.

Software Engineering arXiv cs.AI

An Empirical Study for GUI Test Migration from Android to OpenHarmony System

An empirical study on migrating GUI tests from Android to OpenHarmony, introducing the ITeM-HM approach to improve success rates.

AI/ML arXiv cs.AI

Multi-Agent LLMs Fail to Explore Each Other

Research shows multi-agent LLMs struggle with exploration, leading to the development of MACE, a framework that promotes structured peer selection.

Cybersecurity Hacker News

TS-2026-009: Insecure argument handling in Tailscale SSH permitted root access

A security vulnerability in Tailscale SSH was discovered that allowed root access via insecure argument handling.

AI/ML Hacker News

Solving 20 Erdős Problems with 20 Codex Accounts Running in Parallel

A researcher used 20 Codex accounts in parallel to solve 20 of Erdős's mathematical problems.

Tech Business/VC Hacker News

Data centers have hiked electricity prices on the public by $23B

Data centers have reportedly increased electricity prices for the public by $23 billion.

AI/ML arXiv cs.AI

Flout at Your Own Risk: LLMs Struggle with Pragmatic Cooperativity Under Epistemic Asymmetry

Research explores how LLMs struggle with pragmatic cooperativity when there is an epistemic asymmetry between agents.

AI/ML arXiv cs.AI

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video

A study reveals that Video-LLMs often rely on coarse gender cues rather than actual character tracking in long-form videos.

AI/ML arXiv cs.AI

Controlling Motion Transfer in Diffusion Transformers via Attention Heads

A new framework for controlling motion transfer in Diffusion Transformers (DiTs) by identifying and manipulating specialized attention heads.

AI/ML arXiv cs.AI

AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP

AgentCheck is an open-source workbench for reproducing, intervening, and mitigating failures in tool-using LLM agents over MCP.

AI/ML arXiv cs.AI

The Equilibrium Is the Initialization: Lazy Identity Collapse in Physics-Structured Deep Equilibrium Reasoning

A study warns against 'lazy identity collapse' in physics-structured Deep Equilibrium models, where computation becomes a silent no-op.