All Articles
16760 articles total
SLAPBench: Benchmarking Multimodal Large Language Models for Four-Finger SLAP Fingerprint Verification
Presents SLAPBench, a benchmark for evaluating Multimodal LLMs on four-finger SLAP fingerprint verification.
Recursive Harness Self-Improvement
Introduces Recursive Harness Self-Improvement (RHI) to optimize agent loops and improve execution-trace quality for model training.
Kolmogorov--Arnold Networks for Small Language Models
Evaluates Kolmogorov-Arnold Networks (KANs) as potential replacements for MLPs in small language models, finding limited advantage over strong baselines.
CoWeaver: A Bi-directional, Learnable and Explainable Matching Engine for Mixed Human-Agent Science Collaboration
Proposes CoWeaver, a learnable matching engine designed to facilitate collaboration between human scientists and AI agents.
From Feasibility to Desirability: Plan, Learn, Adapt (PLA) Framework for Personalized On-Device Itinerary Generation
Presents the PLA framework for personalized on-device itinerary generation, combining lightweight planners and a Bradley-Terry reward model.
Evolutionary Algorithm-Guided LLMs for Physics-Informed Neural Network Design
Develops a closed-loop evolutionary algorithm that guides LLMs to design optimal configurations for Physics-Informed Neural Networks (PINNs).
Hard Rules, Soft Preferences: Bridging Reasoning, Learning, and Optimization for Personalized Packing Checklist Generation
Introduces a reasoning-guided learning framework for personalized packing checklists using a symbolic engine and CP-SAT optimizer.
Self-Powered Trailers Promise Leaner Freight Runs
Article discussing self-powered trailers designed to reduce fuel consumption and emissions in freight transport.
Xiaomi-Robotics-1
Discussion regarding Xiaomi's robotics initiatives.
Partial Information Decomposition as a Multi-Contrast 3D MRI Selection Strategy for Resource-Constrained Deep Neural Network Training in Brain Tumor Segmentation
A study on using Partial Information Decomposition to select the most informative MRI input pairs for efficient brain tumor segmentation in deep learning.
AI Trading: Evaluating Large Language Models for Technical Market Analysis
A comparative evaluation of several LLMs, including GPT-4 and FinGPT, on their ability to perform technical market analysis and trading signal generation.
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
Introduction of the Manager Coercion Benchmark to study how AI agents escalate pressure and use deception when managing other AI agents.
Design-Based Supervised Learning with Noisy Human Labels
Proposes Partially Adjudicated Design-Based Supervised Learning (PA-DSL) to correct noisy human labels in automated classifiers.
FLINT: Fingerprinting Federated Learning Architectures from 5G PHY-Layer Side Channels
Presents FLINT, a framework that fingerprints Federated Learning model architectures by analyzing 5G PHY-layer side-channel metadata.
Verbalizable Representations Form a Global Workspace in Language Models
Research identifying a 'global workspace' in LLMs using the Jacobian lens, allowing for the decoding of a model's internal deliberation and strategic thinking.
LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4
A closed-loop AutoML framework utilizing LLMs (GPT-5, GPT-4o, Claude 3.5 Sonnet) to autonomously design neural architectures for cross-lingual OCR.
An Auto-Scaling Approach for Serverless Environments Based on a Multi-Expert Consensus Mechanism
An auto-scaling approach for serverless environments using a multi-expert consensus mechanism and dependency-aware graphs to reduce costs and latency.
LoRA Speedrun – a public wall-clock leaderboard for fine-tuning techniques
A public leaderboard tracking the wall-clock time for various LoRA fine-tuning techniques to optimize training speed.
Harmonizing AI Safety Thresholds
A proposal for harmonizing AI safety capability thresholds across different AI companies to prevent a 'race to the bottom' in safety standards.
CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data
Introduces CRAFT, a method for diagnosing LLM weaknesses using rubric-based evaluations to generate targeted fine-tuning data.