AI/ML arXiv cs.AI

What the Waveform Knows: Transparent-first Speech and Audio Intelligence with Caption Studio

Caption Studio is presented as a transparency-first platform for speech and audio intelligence, utilizing Whisper and pyannote for transcription and diarization.

AI/ML arXiv cs.AI

Strategy-Following Multi-Agent Deep Reinforcement Learning Considering Control Strategies Provided to Other Agents

A study on multi-agent deep reinforcement learning that allows human managers to control specific agents while others adaptively complement the work.

AI/ML arXiv cs.AI

Planning as Emergent Behavior in Reinforcement Learning with Relational Hidden States

Researchers identify that relational hidden-state architectures in model-free RL agents allow for emergent planning behavior, effectively recovering the environment's transition structure.

AI/ML arXiv cs.AI

AutoIndex: Learning Representation Programs for Retrieval

AutoIndex is a framework that optimizes document representation for retrieval systems by searching for executable transformation programs rather than tuning hyperparameters.

AI/ML arXiv cs.AI

Intelligent Multi-UAV Navigation in ITNTNs: A Hierarchical LLM Approach

A hierarchical LLM framework is proposed for UAV navigation, utilizing a cloud-based LLM for global strategy and edge-LLMs for tactical sub-goals to guide a DRL controller.

AI/ML arXiv cs.AI

Mitigating Matthew Effect: Multi-Hypergraph Boosted Multi-Interest Self-Supervised Learning for Conversational Recommendation

HiCore is a novel conversational recommendation framework using multi-hypergraph boosted self-supervised learning to mitigate the 'Matthew effect' of popular item bias.

AI/ML arXiv cs.AI

LatentMT: Machine Translation with Latent Reasoning

LatentMT introduces recurrent computation within hidden states (LoopLMs) for machine translation, allowing small models to achieve performance comparable to much larger ones.

Other arXiv cs.AI

Temporal-Causal Unity as an Operational Framework for Collective Dynamics: Causal-Progress Clocks, Synchronization, and Polarization

The Temporal-Causal Unity (TCU) framework proposes a mathematical model connecting time and causal change to study cognitive and social dynamics, including synchronization and polarization.

Cybersecurity arXiv cs.AI

CPInj: Uncovering Prompt Injection Risks in Textual Collaborative Prompt Optimization

CPInj reveals a critical prompt injection vulnerability in collaborative prompt optimization (TCPO), proposing a defense-oriented aggregation method called APAgg.

AI/ML arXiv cs.AI

Norm or Direction? Decoding Vision Mambas for High-Resolution Vision

Research compares VMamba and MambaOut vision models, finding that VMamba's superior performance in dense prediction tasks comes from how it organizes semantic evidence in token directions.

AI/ML arXiv cs.AI

Deep Learning Estimation of Sex, Age, Height, and Weight from CT-derived Digitally Reconstructed Radiographs

An ensemble of deep learning models (ConvNeXt, ViT, MaxViT) is used to accurately estimate sex, age, height, and weight from CT-derived radiographs.

Cybersecurity arXiv cs.AI

Broken Gates: Re-evaluating Web Bot Defenses in the Age of LLM Agents

A study finds that most web bot defenses (Captchas, etc.) are broadly ineffective against commercial solvers and LLM-based browser agents, emphasizing environment authenticity over behavior.

Other Hacker News

Back to Kagi

A discussion on returning to the Kagi search engine.

Other Hacker News

Businesses with ugly AI menu redesigns

A critique of businesses using AI to redesign their menus, often resulting in poor user experiences.

Other Hacker News

Why Do You Suck at Juggling?

An exploration of the difficulties and mechanics of learning to juggle.

Tech Business/VC TechCrunch

Cascade raises $3.5M to help construction firms find and win projects

Cascade raises $3.5M seed funding to help construction firms find and win projects.

Other TechCrunch

The browser wars aren’t about search anymore — here are the best alternatives to Chrome and Safari

An overview of alternative browsers to Chrome and Safari.

AI/ML arXiv cs.AI

EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration

Introduces EduPanel, a multi-agent LLM judge designed to evaluate the pedagogical quality of teaching videos.

AI/ML arXiv cs.AI

Censoring-Aware In-Context Learning for Generalized Supplier Lead Time Estimation in Supply Chain Planning

Proposes LeadTime-ICL, a censoring-aware in-context learning model for probabilistic supplier lead time forecasting in supply chains.

AI/ML arXiv cs.AI

Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary

Investigates 'operational proto-introspection' in looped language models, testing if models can monitor their own computation quality.