SPIDER Topics

arXiv Research Feed

Fresh ML papers as they land — the research frontier without the firehose. Indexed live by the Starship No. 27 crawl pipeline.

Showing the most recent 30 stories · updated continuously

01
arXiv: Does AI Save Time on Product Design? A Randomized Controlled Experiment of AI Prompt-to-Design Workflows

AI tools for digital product design now offer prompt-to-design capabilities, allowing designers and their non-designer colleagues to create prototypes through conversational workflows with large language models.

arxiv · 2026-09-23
02
arXiv: Diffusion-Induced Spatial Attention Overlapping Community Detection

The paper introduces DISCO, a deep-learning framework that combines structural prior, sparse multi-head attention, and non-negative community-affiliation learning to detect overlapping communities in networks.

arxiv · 2026-09-23
03
arXiv: Automatic depth-based local center clustering via $β$-integrated local depth and adaptive grouping

The paper proposes an automatic clustering method that uses $\beta$-integrated local depth to identify stable exemplars and merge groups based on graph theory.

arxiv · 2026-09-23
04
arXiv: Metrics Failure in LLM-Based Code Vulnerability Repair: An Empirical Study and a Change-Aware Screen

The study examines the reliability of compile rate as a proxy for LLM-based code vulnerability repair, finding it unreliable and dominated by artifacts.

arxiv · 2026-09-23
05
arXiv: EquivSVA: A Formally Verified Dataset of Behavioral Assertions Across Equivalent RTL Implementations

The paper presents EquivSVA, a formally verified dataset of behavioral assertions across equivalent RTL implementations.

arxiv · 2026-09-23
06
arXiv: FleXray: Universal Clinical X-ray Segmentation

FleXray is a generalist model for anatomical segmentation across clinical X-rays, simulating fully-annotated 2D X-rays to train on unseen datasets.

arxiv · 2026-09-23
07
arXiv: Type-Safe Is Not Error-Free: A Constrained Decision Head Follows the Option Name, Not the Rubric Bound to It

Typed decision models guarantee output schema conformity but do not ensure correct interpretation of options.

arxiv · 2026-09-23
08
arXiv: Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents

The paper introduces Growing Harness, a training paradigm that learns an agent harness from a strategy-free scaffold and uses task feedback to localize failures, repair them jointly, and roll back harmful sequences. It aims to move recurring control out of model context into reusable executable code

arxiv · 2026-09-23
09
arXiv: A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem

The paper introduces A2M, a two-stage black-box framework for hijacking MCP agents using semantic matching and execution traces.

arxiv · 2026-09-23
10
arXiv: SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving

The text introduces SWE-Serve as a benchmark for evaluating agents on production inference engineering tasks.

arxiv · 2026-09-23
11
arXiv: CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents

CliffCompaction is an autocompaction technique that reduces cost by up to 50% under a bounded context while maintaining or improving performance on Terminal-Bench and achieving new levels of efficiency for test-time scaling.

arxiv · 2026-09-23
12
arXiv: SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue

The text discusses a dual-track memory system for multi-party dialogue that distinguishes speaker attribution and relational understanding.

arxiv · 2026-09-23
13
arXiv: Agensh: Scaling Organizational Intelligence to 1,024 Agents

The paper introduces Agensh, a scalable self-organized multi-agent harness for complex tasks without a central orchestrator.

arxiv · 2026-09-23
14
arXiv: A Decentralized Partially Observable Team Decision Methodology with Delayed Information Sharing

Decentralized team decision-making in partially observable Markov decision processes with low-rank dynamics and unknown system models.

arxiv · 2026-09-23
15
arXiv: Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs

The paper introduces Flash-dLLM, a training-free inference acceleration framework for fast and memory-efficient diffusion LLMs that addresses I/O bottlenecks in KV-cache-enabled dLLM inference.

arxiv · 2026-09-23
16
arXiv: Blaming Across the Aisle: Political Contrasting and Blame Attribution in the Danish Parliament

This study examines blame attribution in the Danish Parliament using a purpose-built classifier to analyze political discourse over two decades.

arxiv · 2026-09-23
17
arXiv: TransBERT: A Framework for Synthetic Translation in Domain-Specific Language Modeling

The paper introduces TransBERT, a framework for pre-training language models using synthetically translated text.

arxiv · 2026-09-23
18
arXiv: HYDRA: Proactive Android Malware Drift Adaptation via Hierarchical Graph Contrastive Learning

The text discusses a proactive adaptation framework for Android malware detection that learns drift-invariant representations using hierarchical graph contrastive learning.

arxiv · 2026-09-23
19
arXiv: PACT: From Credit Assignment to Critic Alignment

The paper introduces a new framework for understanding token-level credit in reinforcement learning, proving three regularity conditions that uniquely determine it.

arxiv · 2026-09-23
20
arXiv: GitScholar: A Dataset for Predicting AI Research Impact from GitHub Engagement

The paper proposes using GitHub engagement as an additional source to predict AI research impact from GitHub activity.

arxiv · 2026-09-23
21
arXiv: HySparse2: Hybrid Sparse Attention with Two-Level KV Sharing

Hybrid sparse attention architecture with two-level KV sharing for efficient prefill, compact KV-cache storage, and accurate long-context retrieval.

arxiv · 2026-09-23
22
arXiv: Layout-Guided Masking for GROBID: Lightweight Structural Gains in Large-Scale Scientific PDF Ingestion

Layout-Guided Masking for GROBID improves the accuracy of PDF parsing by localizing figure, table, and paratext regions and routing them to specialized models.

arxiv · 2026-09-23
23
arXiv: Learning to Defer with Guidance on Real World Medical Data

Medical image interpretation, learning to defer, real-world medical data.

arxiv · 2026-09-23
24
arXiv: On the Lexical Superstition of Large Language Models for Code Comprehension: Re-evaluation on Code of Low Lexical Quality

Large language models assign disproportionate weight to lexical cues in code renaming, leading to performance degradation when identifier information is removed or misleading.

arxiv · 2026-09-23
25
arXiv: Double Descent and Malign Overfitting in Diffusion Models

Overparameterization leads to catastrophic overfitting in diffusion models, driving them into a memorization regime.

arxiv · 2026-09-23
26
arXiv: Combining Hierarchical Cognitive Process with Process Supervision for Interpretable Scene Safety Understanding

This paper proposes a hierarchical cognitive process model for scene safety understanding, integrating it with process supervision.

arxiv · 2026-09-23
27
arXiv: OMatG-flash: An All-Atom Flow Map with Reinforce Adjoint Matching for Scalable Materials Discovery

The paper introduces OMatG-flash, an all-atom flow map for inorganic crystal structure prediction and de novo generation, which improves inference efficiency and match rates.

arxiv · 2026-09-23
28
arXiv: Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations

The paper investigates harm laundering in GPT models, showing that explicit discriminatory content is transformed rather than removed across generations.

arxiv · 2026-09-18
29
arXiv: RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

Self-Retiring On-Policy Distillation (RetireOPD) for agentic reinforcement learning.

arxiv · 2026-09-18
30
arXiv: PosteriorBench: From Point Estimates to Posterior Matching in Evaluating Generative Inverse Solvers

Generative inverse solvers are evaluated using PosteriorBench to assess their ability to recover the full set of solutions rather than a single best sample.

arxiv · 2026-09-18
Daily round-ups: Tech Digest archive · All facets: Browse topics