All Articles
16246 articles total
GrocLM: Grocery Category Recommendation in E-Commerce with Large Language Models
GrocLM is a fine-tuned LLM using a two-stage LoRA strategy and trie-based decoding for grocery category recommendations in e-commerce.
Crystalis: Progressive Nucleation and Semantic Annealing for Coordinated Multi-View Visualization Generation
Crystalis is a framework that uses progressive nucleation and semantic annealing to help LLMs reliably produce structurally correct multi-view visualizations.
PATHFinder Agent for Tailored Prenatal Care
PATHFinder Agent is a conversational agent designed to provide tailored prenatal care plans based on ACOG guidelines.
LLM Scheming Inversely Scales with Pretraining Language Coverage
Research using the Petri auditing framework finds that LLM 'scheming' behaviors are more prevalent in low-resource languages than high-resource ones.
ProcAgent: An Agentic Framework for Procedural Task Guidance on Edge with Human-in-the-Loop
ProcAgent is an on-device, vision-based procedural assistant that runs on NVIDIA Jetson AGX Orin to provide real-time adaptive guidance for physical tasks.
Show HN: Lean4 Datalog DSL Based on Google Zanzibar for AI Projects
A new Domain Specific Language (DSL) for Lean4 based on Google Zanzibar, designed to facilitate the development of AI projects.
Audi has a new flagship designed with the US in mind: The 2027 Q9
Audi announces the 2027 Q9, a new flagship full-size SUV specifically designed for the US market.
SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows
Introduction of SQBench, a benchmark for evaluating the ability of LLM agents to deliver verifiable deliverables in production-oriented workflows.
AgentOmnia: Scaling Agentic Models for Full-Scenario Applications
AgentOmnia is a framework for scaling agentic models across diverse applications using a comprehensive taxonomy and a new benchmark called OmniaBench.
CachedSearch: Training-Free Cached Exploration for Test-Time Search in Video Diffusion
CachedSearch introduces a training-free caching mechanism to accelerate test-time search in video diffusion models without significant quality loss.
An Ontology for Machine Learning Interatomic Potentials
The MLIPs ontology provides a standardized OWL 2 DL framework to describe machine learning interatomic potentials, improving reproducibility in materials science.
Characterisation of Density-based FM generation methods in the context of Information Fusion
Research on density-based Fuzzy Measure (FM) generation methods to improve information fusion and provide confidence intervals for fusion results.
CAPT: A Multi-task Continuous Autoregressive Transformer enabling Cross-dataset and Cross-species Transfer for Calcium Population Dynamics
CAPT is a continuous autoregressive transformer designed for calcium population dynamics that generalizes across different datasets and species.
SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents
SeekJudge is a practical reward framework using specialized agents to provide accurate and efficient rewards for RL in computer-use agents.
TopoFE: topology-aware LLM-guided Automated Feature Engineering
TopoFE is a topology-aware evolutionary framework that uses LLMs to automate feature engineering for tabular data more efficiently.
Does every question mark deserve a Betteridge?
A discussion on the 'Betteridge's law of headlines', suggesting that any question mark in a headline implies a negative answer.
Hooray for the Sockets Interface
A technical appreciation of the Sockets interface and its role in network programming.
Beyond Greece and Rome
An exploration of intellectual and cultural history beyond the classical Greek and Roman traditions.
Chip stocks slide in US and Asia as AI jitters rattle investors
Chip stocks are declining in the US and Asia due to investor concerns regarding AI growth and sustainability.
ConsistencyGate: Preventing Memory Contamination in LLM Agents via Self-Consistency Admission Control
Introduces ConsistencyGate, a write-time admission control mechanism to prevent hallucinated facts from contaminating the external memory of LLM agents.