AI/ML arXiv cs.AI

SLAPBench: Benchmarking Multimodal Large Language Models for Four-Finger SLAP Fingerprint Verification

Presents SLAPBench, a benchmark for evaluating Multimodal LLMs on four-finger SLAP fingerprint verification.

AI/ML arXiv cs.AI

Recursive Harness Self-Improvement

Introduces Recursive Harness Self-Improvement (RHI) to optimize agent loops and improve execution-trace quality for model training.

AI/ML arXiv cs.AI

Kolmogorov--Arnold Networks for Small Language Models

Evaluates Kolmogorov-Arnold Networks (KANs) as potential replacements for MLPs in small language models, finding limited advantage over strong baselines.

AI/ML arXiv cs.AI

CoWeaver: A Bi-directional, Learnable and Explainable Matching Engine for Mixed Human-Agent Science Collaboration

Proposes CoWeaver, a learnable matching engine designed to facilitate collaboration between human scientists and AI agents.

AI/ML arXiv cs.AI

From Feasibility to Desirability: Plan, Learn, Adapt (PLA) Framework for Personalized On-Device Itinerary Generation

Presents the PLA framework for personalized on-device itinerary generation, combining lightweight planners and a Bradley-Terry reward model.

AI/ML arXiv cs.AI

Evolutionary Algorithm-Guided LLMs for Physics-Informed Neural Network Design

Develops a closed-loop evolutionary algorithm that guides LLMs to design optimal configurations for Physics-Informed Neural Networks (PINNs).

AI/ML arXiv cs.AI

Hard Rules, Soft Preferences: Bridging Reasoning, Learning, and Optimization for Personalized Packing Checklist Generation

Introduces a reasoning-guided learning framework for personalized packing checklists using a symbolic engine and CP-SAT optimizer.

Other Hacker News

Self-Powered Trailers Promise Leaner Freight Runs

Article discussing self-powered trailers designed to reduce fuel consumption and emissions in freight transport.

Hardware/Chips Hacker News

Xiaomi-Robotics-1

Discussion regarding Xiaomi's robotics initiatives.

AI/ML arXiv cs.AI

Partial Information Decomposition as a Multi-Contrast 3D MRI Selection Strategy for Resource-Constrained Deep Neural Network Training in Brain Tumor Segmentation

A study on using Partial Information Decomposition to select the most informative MRI input pairs for efficient brain tumor segmentation in deep learning.

AI/ML arXiv cs.AI

AI Trading: Evaluating Large Language Models for Technical Market Analysis

A comparative evaluation of several LLMs, including GPT-4 and FinGPT, on their ability to perform technical market analysis and trading signal generation.

AI/ML arXiv cs.AI

Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation

Introduction of the Manager Coercion Benchmark to study how AI agents escalate pressure and use deception when managing other AI agents.

AI/ML arXiv cs.AI

Design-Based Supervised Learning with Noisy Human Labels

Proposes Partially Adjudicated Design-Based Supervised Learning (PA-DSL) to correct noisy human labels in automated classifiers.

Cybersecurity arXiv cs.AI

FLINT: Fingerprinting Federated Learning Architectures from 5G PHY-Layer Side Channels

Presents FLINT, a framework that fingerprints Federated Learning model architectures by analyzing 5G PHY-layer side-channel metadata.

AI/ML arXiv cs.AI

Verbalizable Representations Form a Global Workspace in Language Models

Research identifying a 'global workspace' in LLMs using the Jacobian lens, allowing for the decoding of a model's internal deliberation and strategic thinking.

AI/ML arXiv cs.AI

LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4

A closed-loop AutoML framework utilizing LLMs (GPT-5, GPT-4o, Claude 3.5 Sonnet) to autonomously design neural architectures for cross-lingual OCR.

Software Engineering arXiv cs.AI

An Auto-Scaling Approach for Serverless Environments Based on a Multi-Expert Consensus Mechanism

An auto-scaling approach for serverless environments using a multi-expert consensus mechanism and dependency-aware graphs to reduce costs and latency.

AI/ML Hacker News

LoRA Speedrun – a public wall-clock leaderboard for fine-tuning techniques

A public leaderboard tracking the wall-clock time for various LoRA fine-tuning techniques to optimize training speed.

AI/ML arXiv cs.AI

Harmonizing AI Safety Thresholds

A proposal for harmonizing AI safety capability thresholds across different AI companies to prevent a 'race to the bottom' in safety standards.

AI/ML arXiv cs.AI

CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data

Introduces CRAFT, a method for diagnosing LLM weaknesses using rubric-based evaluations to generate targeted fine-tuning data.