AI/ML arXiv cs.AI

StructureClaw: Traceable LLM Agents and an Executable Benchmark for Structural Engineering Workflows

Introduction of StructureClaw, a workbench and executable benchmark for evaluating LLM agents in complex structural engineering workflows.

Hardware/Chips Hacker News

Show HN: Simulator for a custom 8-bit discreet logic computer

A user shared a simulator for a custom 8-bit computer built using discrete logic, appealing to those interested in low-level hardware emulation.

Cybersecurity TechCrunch

FBI arrests man accused of using Steam games to drain victims’ crypto wallets

The FBI arrested a student who used fake Steam games to distribute malware and steal cryptocurrency from victims.

Other TechCrunch

Parents want safer phones for kids. These companies are answering the call.

New companies are launching phones designed specifically for children, offering limited features and minimalist designs for safety.

Other The Verge

Is America ready for this quirky Jeep-looking EV that can park itself?

Chip Motors is launching a small, boxy electric vehicle in Miami, reflecting a niche trend toward microcars in the US.

Other Ars Technica

Fubo hikes prices by $15 after restoring some NBCU channels lost in November

Fubo is increasing its subscription prices by $15 following the restoration of some NBCUniversal channels.

Cybersecurity Ars Technica

San Francisco orders Apple, Google to remove nudify apps from app stores

San Francisco is directing Apple and Google to remove 'nudify' apps from their app stores due to safety concerns.

Other Ars Technica

Ars is looking for a senior technology reporter, and you might be it!

Ars Technica is hiring a senior technology reporter to cover hardware, CPUs, GPUs, and NAS.

AI/ML arXiv cs.AI

Large Language Models for Code Generation from Multilingual Prompts: A Curated Benchmark and a Study on Code Quality

Researchers introduced a multilingual benchmark to study how prompt language affects the quality and correctness of code generated by LLMs.

AI/ML arXiv cs.AI

Evaluating Epistemic Uncertainty: Beyond OOD Detection and Active Learning

This paper proposes a new theoretical framework for evaluating epistemic uncertainty in AI, moving beyond OOD detection and active learning.

AI/ML arXiv cs.AI

Can LLMs Build a MaxSAT Solver from Papers? The CoreForge Experience

The CoreForge project demonstrates using LLMs to implement a MaxSAT solver based on research papers rather than existing codebases.

Other Hacker News

Show HN: Explore the Workspaces of Modern Creators

A 'Show HN' post showcasing a gallery of workspaces used by modern creators.

AI/ML Hacker News

Kimi K3, and what we can still learn from the pelican benchmark

Analysis of the Kimi K3 model and insights derived from the pelican benchmark.

Hardware/Chips The Verge

Samsung’s redesigned Z Fold 8 with a wide display just leaked

Leaked images and specs of the Samsung Galaxy Z Fold 8, featuring a wider display and Snapdragon 8 Elite processor.

AI/ML arXiv cs.AI

Harnessing LLMs for Reliable Academic Supervision: A Comparative Study

A study on 'harness engineering' to make LLMs reliable for academic supervision, demonstrating that structured scaffolding can outperform larger base models.

AI/ML arXiv cs.AI

VideoSEMA: a scalable and efficient Mamba-like attention for video understanding

Introduction of VideoSEMA, a scalable attention model for video understanding that combines Mamba-like spatial attention and temporal softmax attention.

AI/ML arXiv cs.AI

The Misclassification of Autistic Writing as AI-Generated

Research indicating that AI detection models may be biased against autistic writers, frequently misclassifying their writing as AI-generated.

AI/ML arXiv cs.AI

FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models

FoMoVLA is a framework that improves Vision-Language-Action models by jointly learning future feature foresight and sparse 2D point tracking.

AI/ML arXiv cs.AI

Large Audio Language Models for Spoofing-Aware Speaker Verification

Evaluation of Large Audio Language Models (LALMs) for spoofing-aware speaker verification, finding that task-specific adaptation is necessary for effectiveness.

AI/ML arXiv cs.AI

Dialogue Summarization with Emotion Dynamics Using Topic- and Participant-Centric Decomposition

A proposed dialogue summarization framework that models semantic and emotion dynamics using a hierarchical Chain-of-Agents approach.