All Articles
17277 articles total
InvestPhilBench: A Multi-Layer Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy
Introduction of InvestPhilBench, a multi-layer benchmark to evaluate LLMs on their ability to apply expert investment philosophy procedural reasoning.
Phia accused of ‘cookie stuffing,’ taking affiliate credit on purchases it didn’t earn
Shopping startup Phia, founded by Phoebe Gates, is accused of 'cookie stuffing' to fraudulently gain affiliate commissions on sales it did not earn.
OpenCoF: Learning to Reason Through Video Generation
Researchers introduce OpenCoF, a framework and dataset for enhancing reasoning capabilities in video generation models through 'Chain-of-Frame' reasoning.
IFAR: Multi-Perspective and Multi-Level Causal Discovery with LLMs
The IFAR framework is proposed to improve multi-perspective and multi-level abductive reasoning in LLMs, showing significant F1 score improvements over existing methods.
Goal-Driven Reasoning in DatalogMTL with Magic Sets
A new reasoning method for DatalogMTL using magic sets technique is introduced to reduce computational complexity in temporal reasoning tasks.
Dual-Difficulty Curriculum Learning for Direct Preference Optimization
GSP-Curri-DPO introduces a group-wise self-paced learning framework for LLM alignment, optimizing the learning trajectory based on prompt complexity and pairwise distinguishability.
A Vision Toward Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents
A vision paper proposing a shift from massive general-purpose LLMs to energy-efficient, domain-specific multimodal agents for critical application domains.
MetaHGNIE: Meta-Path Induced Hypergraph Contrastive Learning in Heterogeneous Knowledge Graphs
DualHNIE is a dual-channel hypergraph learning framework designed for better node importance estimation in heterogeneous knowledge graphs.
SimRPD: Optimizing Recruitment Proactive Dialogue Agents through Simulator-Based Data Evaluation and Selection
SimRPD is a three-stage framework that uses a high-fidelity user simulator to optimize training data for recruitment-focused proactive dialogue agents.
Conversational AI for Rapid Scientific Prototyping: A Case Study on ESA's ELOPE Competition
A case study on using ChatGPT for rapid scientific prototyping in a lunar trajectory estimation competition, highlighting both efficiency gains and reliability issues.
RetailBench: Evaluating Long-Horizon Autonomous Decision-Making and Strategy Stability of LLM Agents in Realistic Retail Environments
RetailBench is introduced as a simulation benchmark to evaluate the long-horizon autonomous decision-making and strategy stability of LLM agents in retail environments.
Preemption is GC for memory reordering (2019)
A discussion on how preemption in operating systems can be viewed as a form of garbage collection for memory reordering.
Meta removes controversial AI feature on Instagram after backlash
Meta has removed a controversial AI feature on Instagram following significant user backlash.
No, Flock isn’t threatening people for debating surveillance
Flock Safety is facing criticism for allegedly attempting to silence discussions about its surveillance technology via cease and desist letters.
Meta turns off the Instagram feature that let users make AI deepfakes of public accounts
Meta disables an Instagram AI feature that allowed users to create deepfakes of public accounts without permission.
A Practical Investigation of Training-free Relaxed Speculative Decoding
An investigation into training-free relaxed speculative decoding for accelerating LLM sampling, highlighting capability-speed trade-offs.
ProjAgent: Procedural Similarity Retrieval for Repository-Level Code Generation
Introduction of ProjAgent, a system for repository-level code generation using procedural similarity retrieval to better handle complex dependencies.
Pose-to-Biomechanics: Bridging 3D Human Pose Estimation and Biomechanical Attribute Prediction
BioModule is proposed as a lightweight plug-in for 3D pose estimators to predict biomechanical attributes for motion analysis.
Validity of LLMs as data annotators: AMALIA on authority
A study on the AMALIA-9B model's validity as a data annotator for European Portuguese, questioning its reliance on surface correlates.
Dimensionality Reduction Meets Network Science: Sensemaking on UMAP's kNN Graph
Research demonstrating how applying graph algorithms to UMAP's internal kNN graph can enhance high-dimensional data sensemaking.