AI/ML arXiv cs.AI

Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not Ground Truth

This research challenges the use of human agreement as the gold standard for LLM qualitative coding, showing that blind expert verification can favor LLM interpretations over human consensus.

AI/ML arXiv cs.AI

TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models

TORUS is introduced as a self-coherence test for unified audio models to determine if their audio generation and understanding capabilities are aligned.

Other arXiv cs.AI

Design Concept: Scaffolding Geopolitical Reflection Among Tech Workers

A speculative HCI design proposal for an AI-enabled narrative system aimed at encouraging geopolitical reflection among technology workers.

AI/ML arXiv cs.AI

Gated Q-learning: Add Off-Policy Bias to Taste

Gated Q-learning is proposed as a framework to manage off-policy bias in reinforcement learning by interpolating between eliminating bias and ignoring it via a gating mechanism.

AI/ML arXiv cs.AI

FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation

FairFund-Bench is a new benchmark for evaluating distributive bias in LLM resource allocation, revealing that audit formats significantly influence bias detection.

Cybersecurity arXiv cs.AI

DiffAttack: Evasion Attacks Against Face Recognition via Latent Diffusion Models

DiffAttack uses latent diffusion models to create high-quality adversarial images that can evade facial recognition systems with high success rates.

Cybersecurity arXiv cs.AI

Retrieval-Driven Training-Free AI-Generated Video Attribution

A training-free attribution paradigm is introduced to identify the generative source of AI-generated videos using a generative fingerprint-based pipeline.

Other Hacker News

Less Coffee, Better Sleep

A discussion on lifestyle choices regarding caffeine consumption and its impact on sleep quality.

Cybersecurity Hacker News

What DMARC Protects You From, and What It Does Not

An explanation of the DMARC protocol and its specific capabilities and limitations in email security.

Other Hacker News

MPs demand answers on Fujitsu's inclusion in lucrative frameworks

Political scrutiny in the UK regarding Fujitsu's inclusion in government frameworks.

Cybersecurity Hacker News

PISIGuard: Protect your personal and sensitive info when you chat with AI

A project called PISIGuard designed to protect sensitive information during AI interactions.

Software Engineering Hacker News

The true power of regular expressions (2012)

A technical retrospective on the power and utility of regular expressions.

Homelab/Self-Hosting Hacker News

Show HN: We Fixed UniFi's Slow PPPoE Performance with PPPoE Half-Bridge

A developer-focused fix for slow UniFi PPPoE performance using a half-bridge method.

Tech Business/VC TechCrunch

A Marc Benioff-backed startup thinks AI can solve the AI deployment problem

A startup named June emerged from stealth with significant funding to simplify AI deployment.

AI/ML arXiv cs.AI

A Unified Benchmark of Deep Learning Models for Multi-task 3D Brain Tumor Segmentation from Magnetic Resonance Imaging

A research paper providing a unified benchmark for deep learning models used in 3D brain tumor segmentation.

AI/ML arXiv cs.AI

TextCloak: Thwarting Unauthorized LLM Exploitation via RL-Driven Unlearnable Text

A research paper proposing TextCloak, an RL-driven framework to prevent unauthorized LLM exploitation using unlearnable text.

Software Engineering arXiv cs.AI

Validation Evidence in LLM Repair Agents: How Much of What Passes Actually Tests the Bug?

A study analyzing the effectiveness of validation evidence in LLM-driven software repair agents.

Software Engineering Hacker News

Octane – React's programming model, compiled

An exploration of Octane, a project that compiles React's programming model for potentially better performance.

AI/ML arXiv cs.AI

WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization

Introduces WitCert, a runtime observability tool for KV-cache quantization in LLMs that provides provable bounds on compression damage.

AI/ML arXiv cs.AI

A user's guide to PINNs in geometric analysis: lessons from the asymptotic Plateau problem

A guide on using Physics-Informed Neural Networks (PINNs) for geometric analysis, detailing optimizations that reduce training costs by 40-50x.