AI/ML arXiv cs.AI

Understanding User Experiences of Computer Use Agents: Design Space and Opportunities for Building Agent UX Prototypes

Analyzes the UX design space for computer use agents and introduces AgentUXlab, a tool for prototyping and evaluating agent user experiences.

AI/ML arXiv cs.AI

Long-Term PM2.5 Forecasting Using a DTW-Enhanced CNN-GRU Model

Proposes a DTW-enhanced CNN-GRU model for stable long-term PM2.5 air quality forecasting in resource-constrained urban environments.

AI/ML arXiv cs.AI

DeepVRegulome: DNABERT-based deep-learning framework for predicting the functional impact of short genomic variants on the human regulome

Introduces DeepVRegulome, a framework using fine-tuned DNABERT models to predict the functional impact of genomic variants on the human regulome.

AI/ML arXiv cs.AI

RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension

Introduces RefBench-PRO, a benchmark for referring expression comprehension, and Ref-R1, an RL-based learning scheme to improve localization accuracy.

AI/ML arXiv cs.AI

Deep Delta Learning

Presents Deep Delta Learning (DDL), a structured residual update for Transformers that enables targeted edits to the residual state, improving language modeling quality.

AI/ML arXiv cs.AI

Measuring the State of Open Science in Transportation Using Large Language Models

Uses LLMs to automatically measure open science practices (code and data sharing) in transportation research, revealing low adoption rates.

AI/ML arXiv cs.AI

Picasso: Holistic Scene Reconstruction with Physics-Constrained Sampling

Introduces Picasso, a physics-constrained scene reconstruction pipeline and dataset that ensures geometrically and physically plausible multi-object reconstructions.

Software Engineering Hacker News

Kuna: Decompiler Development in the Age of Coding Agents

Discussion on the development of the Kuna decompiler and how coding agents are influencing the field of reverse engineering.

Other Hacker News

Thanatos Rising

A Hacker News thread titled 'Thanatos Rising', though the snippet provides no substantive content for a technical analysis.

AI/ML arXiv cs.AI

Representation Capacity-Matched QNN-SNN Twin Construction for Rate-Encoded SNNs

A research paper proposing a capacity-matched construction between QNNs and SNNs to provide a fair energy efficiency comparison for neuromorphic hardware.

AI/ML arXiv cs.AI

A context-adaptive policy framework for robust and reactive robotic manipulation via uncertainty-aware imitation learning

Presents a context-adaptive policy framework for robotic manipulation using uncertainty-aware imitation learning and a Mixture of Experts (MoE) formulation.

AI/ML arXiv cs.AI

COMPOL: A Unified Neural Operator Framework for Scalable Multi-Physics Simulations

Introduces COMPOL, a neural operator framework designed to improve the scalability and accuracy of multi-physics simulations using attention-based aggregation.

AI/ML arXiv cs.AI

Localizing Persona Representations in LLMs

A study analyzing where personas are encoded in LLM representation spaces, finding that divergence primarily occurs in the final third of decoder layers.

AI/ML arXiv cs.AI

Towards Understanding the Cognitive Habits of Large Reasoning Models

Introduces CogTest, a benchmark to evaluate human-like cognitive habits in Large Reasoning Models (LRMs) through their Chain of Thought (CoT) patterns.

AI/ML arXiv cs.AI

TaylorPODA: A Taylor Expansion-Based Method to Improve Post-Hoc Attributions for Opaque Models

Proposes TaylorPODA, a model-agnostic local attribution method based on Taylor expansion to improve the explainability of opaque AI models.

AI/ML arXiv cs.AI

Fairness Is Not Enough: Auditing Competence and Intersectional Bias in AI-powered Resume Screening

An audit of AI resume screening tools finding that perceived fairness often stems from a lack of evaluative competence rather than actual neutrality.

AI/ML arXiv cs.AI

Annotation-Assisted Learning of Treatment Policies From Multimodal Electronic Health Records

Introduces AACE, an annotation-assisted approach to causal policy learning for multimodal electronic health records to improve medical treatment decisions.

Other Hacker News

NSF pilots 4-year PhDs with industry research placements

The NSF is piloting a program to implement 4-year PhDs that include industry research placements to better align academic research with industry needs.

Software Engineering Hacker News

Logic for Programmers by Hillel Wayne

Logic for Programmers is a resource by Hillel Wayne focusing on applying formal logic to software development.

AI/ML arXiv cs.AI

Sheet As Token: A Graph-Enhanced Representation for Multi-Sheet Spreadsheet Understanding

The Sheet As Token (SAT) framework improves multi-sheet spreadsheet understanding for LLMs by treating worksheets as unified semantic units and using a graph-enhanced representation.