AI/ML arXiv cs.AI

Ensemble Learning for Large Language Models in Text and Code Generation: A Survey

A comprehensive survey on ensemble learning techniques for LLMs in text and code generation, categorizing methods like weight merging and MoE.

AI/ML arXiv cs.AI

Multimedia and Visual Analytics in the Agentic Era

Proposes a framework integrating multimedia and visual analytics to support professional users in the agentic era of AI.

Cybersecurity arXiv cs.AI

MuTRAP: Multi-trigger Trojans Attacking Robot Task Planning Systems

Introduces MuTRAP, a multi-trigger trojan attack targeting LLM-assisted robot task planning systems to highlight security vulnerabilities.

AI/ML arXiv cs.AI

Minimisation of Quasar-Convex Functions Using Random Zeroth-Order Oracles

Analyzes the minimization of quasar-convex functions using random zeroth-order oracles with theoretical convergence guarantees.

Tech Business/VC Hacker News

Zombie unicorns are haunting Silicon Valley

A discussion on 'zombie unicorns' in Silicon Valley, referring to highly valued startups that are no longer growing but remain afloat.

AI/ML arXiv cs.AI

TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation

Introduction of TouchThinker, a framework and million-scale dataset designed to scale tactile commonsense reasoning for embodied AI agents.

AI/ML arXiv cs.AI

Repeated Shared Access Enables Grokking, but Edit Propagation Depends on an Addressable Memory

Research on factual edit propagation in LLMs, finding that addressable memory is the primary driver for successful edit propagation rather than loop recurrence.

AI/ML arXiv cs.AI

When Preferences Fail to Become Incentives: A Utility-Behavior Gap in Large Language Models

A study demonstrating a 'utility-behavior gap' in LLMs, where reported preferences in choice paradigms do not translate into behavioral incentives in real-world tasks.

AI/ML arXiv cs.AI

IPO Finance Agent: Evaluation of LLM Financial Analysts beyond Finance Agent v2, with Automated Rubric Generation -- the Case of the SpaceX (SPCX) IPO

Presentation of the IPO Finance Agent benchmark, which uses contextual retrieval and automated rubric generation to evaluate LLM financial analysis of S-1 filings.

AI/ML arXiv cs.AI

HOLMES: Evaluating Higher-Order Logical Reasoning in LLMs

Introduction of HOLMES, a benchmark for evaluating higher-order logical reasoning in LLMs across law and finance domains.

AI/ML arXiv cs.AI

Invariant Graph Representations for Continuous-Time Dynamic Graphs Under Distribution Shifts

Proposed CIR, a framework for learning invariant graph representations for continuous-time dynamic graphs to improve robustness under distribution shifts.

AI/ML arXiv cs.AI

When AI Meets Finance (StockAgent): Large Language Model-based Stock Trading in Simulated Real-world Environments

Development of StockAgent, a multi-agent LLM system designed to simulate real-world stock trading behaviors and analyze the impact of external factors.

AI/ML arXiv cs.AI

CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark

Introduction of CORE-Bench, a benchmark to measure the ability of AI agents to computationally reproduce scientific research results.

AI/ML arXiv cs.AI

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning

A new Bayesian approach called Variational Model Merging to improve the estimation of Pareto fronts in multitask finetuning for transformers.

Software Engineering Hacker News

Cloudflare launched self-managed OAuth for all

Cloudflare has released a self-managed OAuth service, allowing developers to handle authentication more easily across their own infrastructure.

Software Engineering Hacker News

Show HN: Write SaaS apps where users control where their data is stored

A new SaaS approach allows users to maintain control over where their data is stored, moving away from centralized data ownership.

Software Engineering Hacker News

Ask HN: Where is our profession (programmer) going?

A community discussion on Hacker News regarding the future trajectory and evolution of the programming profession.

AI/ML arXiv cs.AI

Grounded Chess Reasoning in Language Models via Master Distillation

Researchers introduce 'Master Distillation', a framework to distill expert system reasoning into compact LLMs for specialized domains like chess.

AI/ML arXiv cs.AI

Subjective-Graph LLM Agents for Simulating Uncertainty in Classroom Social Perception

A study on simulating social perception and uncertainty in classrooms using subjective-graph LLM agents and Bayesian fusion.

AI/ML arXiv cs.AI

Riemann-Bench: A Benchmark for Moonshot Mathematics

Riemann-Bench is a private benchmark designed to evaluate AI's ability to solve research-level mathematics, highlighting a gap in current model capabilities.