AI/ML arXiv cs.AI

EComAgentBench: Benchmarking Shopping Agents on Long-Horizon Tasks with Distributed Hidden Intent

EComAgentBench introduces a new benchmark for evaluating LLM-based shopping agents on long-horizon tasks with hidden user intent.

AI/ML arXiv cs.AI

LongWebBench: Evaluating Structural and Functional Webpage Generation in Long-Horizon Settings

LongWebBench provides a framework and benchmark for evaluating structural and functional webpage generation in long-horizon VLM settings.

AI/ML arXiv cs.AI

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs

The E3RL framework introduces erasable reinforcement learning to allow LLMs to self-heal reasoning errors in long-sequence logical tasks.

AI/ML arXiv cs.AI

DecoSearch: Complexity-Aware Routing and Plan-Level Repair for Text-to-SQL

DecoSearch is a training-free framework for Text-to-SQL that uses complexity-aware routing and plan-level repair to handle complex queries.

Other Hacker News

Why do commercial spaces sit vacant?

A discussion on the economic and structural reasons why commercial real estate remains vacant.

Other Hacker News

Hacker News but for Independent Blogs

A project or proposal for a discovery platform focused specifically on independent blogs, similar to the Hacker News model.

Other Hacker News

Making 'food out of thin air' (2024)

An exploration of technology used to synthesize food from atmospheric elements.

Other Hacker News

From Chesterton's fence to Chesterton's gap

A philosophical exploration of 'Chesterton's fence' and the concept of 'Chesterton's gap'.

AI/ML arXiv cs.AI

Brick-DICL: Dynamic In-Context Learning for Automated Brick Schema Classification

Introduces Brick-DICL, a dynamic in-context learning framework to automate the classification of Building Management System points into the Brick schema using RAG.

AI/ML arXiv cs.AI

FinAcumen: Financial Multimodal Reasoning via Self-Evolving Experience Memory Harness

Presents FinAcumen, a financial reasoning agent that uses a selective experience memory bank to improve multimodal reasoning and tool routing.

AI/ML arXiv cs.AI

Beyond Domains: Reusing Web Skills via Transferable Interaction Patterns

Introduces SkillMigrator, an agent that learns transferable interaction patterns (TIPs) to reuse web skills across different sites based on layout similarity.

AI/ML arXiv cs.AI

From Brewing to Resolution: Tracing the Internal Lifecycle of Code Reasoning in LLMs

Investigates the internal 'brewing' and 'resolution' lifecycle of code reasoning in LLMs to explain why models fail on specific coding tasks.

AI/ML arXiv cs.AI

Using Cognitive Models to Improve Language Model Simulation of Human Persuasion Games

Proposes Equation-to-Behavior Prompting and RL to make LLMs better simulate human decision-making and biases in persuasion games.

AI/ML arXiv cs.AI

FllumaOne: A Code-Native Multimodal CAD Dataset with Executable Programs and Kernel-Validated Feature Histories

Introduces FllumaOne, a code-native multimodal CAD dataset consisting of 100,000 samples generated via executable Python programs in the Flluma CAD system.

Tech Business/VC Hacker News

The founder's playbook: Building an AI-native startup

A discussion on the strategy and playbook for building startups focused on AI-native architectures.

Other Hacker News

Subterranean fungi networks more than 100 quadrillion km in length

A scientific report regarding the massive scale of subterranean fungi networks.

Hardware/Chips Hacker News

Chameleon Ultra: a flashdrive sized NFC toolkit

Introduction of Chameleon Ultra, a compact NFC toolkit designed for security testing and tool-like functionality.

Other Hacker News

Semiclassical Gravity Efficiently Solves NP-Complete Problems

A theoretical paper exploring how semiclassical gravity might be used to efficiently solve NP-complete problems.

Other Hacker News

Lattice Triangles Are Rare

A mathematical finding demonstrating that lattice triangles are rare.

AI/ML arXiv cs.AI

LLM-as-Judge in Education: A Curriculum-Grounded Marking Pipeline

Proposes a curriculum-grounded pipeline that uses LLMs as judges for marking university admission exams.