All Articles
17188 articles total
Task-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026
A task-specific two-agent architecture utilizing confidence calibration and incremental reasoning achieved top results in the QANTA 2026 multimodal question answering challenge.
The console wars have been lost
A Hacker News discussion regarding the perceived end of competitive console gaming eras.
This free Mac app reveals the truth about your mystery USB-C cables
A free macOS utility called WhatCable that identifies USB-C cable capabilities using Apple Silicon data.
On-Device Adaptive Battery Power Prediction for Electric Vehicles
Research on on-device adaptive power prediction models for electric vehicles to handle distribution shifts.
Self-Guided Test-Time Training for Long-Context LLMs
Introduction of Self-Guided Test-Time Training (S-TTT) to improve long-context reasoning in LLMs.
SVF-CR: Synchronized Visual-Facial Cross-Refinement for Multimodal Ambivalence and Hesitancy Recognition
A multimodal framework (SVF-CR) for recognizing subtle behavioral states like ambivalence and hesitancy.
A Sovereign, Open-Source Foundation Model for German and English
Release of Soofi S 30B-A3B, an open-source MoE Mamba Transformer model optimized for German and English.
Test-Time Scaling for Small VLMs on Multilingual Visual MCQ
An analysis of how test-time scaling impacts small vision-language models on multilingual benchmarks.
Parameter-Efficient Vision-Language Adaptation with Continuous Metadata Conditioning for Animal Re-Identification
A parameter-efficient CLIP adaptation framework designed for long-term animal re-identification.
Practical Source Code Recovery from Binary Functions Using Anchor-Based Retrieval and LLM Reasoning
A pipeline for recovering source code from stripped binary functions using LLM reasoning and anchor retrieval.
Decoupling Language Guidance from Backbones for Text-Guided Medical Segmentation
A backbone-transferable hierarchical adapter framework (BTHA) for text-guided medical image segmentation.
Social media limits are coming for teens across Europe
The European Union is considering strict new legislation to limit social media access for children and teenagers, potentially including age bans and phased access.
Letting the Data Speak: Extracting Keywords from Crowdsourced Collections with AI
Researchers evaluated various NLP techniques for automated keyword extraction in crowdsourced collections, finding that open-weight extractive models are most suitable for responsible deployment.
WILDTRACE: Benchmarking Natural Evidence Trails in Long-Context Reasoning
The WILDTRACE benchmark is introduced to evaluate long-context reasoning in LLMs using naturally occurring evidence trails in complex documents.
Shortcut Trajectory Planning for Efficient Offline Reinforcement Learning
Shortcut Trajectory Planning (STP) is proposed as an efficient offline reinforcement learning framework that reduces inference costs and simplifies the training pipeline.
Deceptive Grounding: Entity Attribution Failure in Clinical Retrieval-Augmented Generation
The paper identifies 'deceptive grounding' in clinical RAG systems, where evidence is attributed to the wrong entity despite appearing factually grounded.
CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation
CtrlVTON and VIP-SAM are introduced to provide controllable virtual try-on capabilities through visual-instance-prompt segmentation.
Diversifying to Verify: When Task-Equivalent Programs Differ in Verifiability
Diversify2Verify explores how implementation diversity helps LLMs find more easily verifiable program variants using the Why3 verification framework.
When Routes Run Out: Adversarial Co-Learning and Explainable Robustness in Quantum Repeater Networks
The authors study adversarial co-learning in quantum repeater networks and provide an open-source explanation workflow for these network games.
STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU
STEEL is an open-source FlashAttention implementation for AMD's XDNA NPU that significantly reduces energy consumption and latency for long-sequence inference.