Research Reality Cards
Dense arXiv abstracts distilled into 3-bullet builder takeaways. No hype, no padding.
I-CARE: Analysis of interference-related phenomena in a controllable, diverse and representative unlearning setting for text-to-image models
Not provided in the content
I-CARE formalizes interference as a first-class object of study in generative unlearning, allowing for systematic and reproducible analysis across various settings.
The methodology enables meaningful analysis of interference patterns across multiple unlearning settings.
The paper does not specify limitations or reproducibility concerns explicitly.
Incremental Risk Assessment of Progressive Elder Financial Scams via Instruction-Tuned Small Language Models
Author1, Author2, Author3 +2 more
The study demonstrates that fine-tuned small language models can effectively perform turn-aware risk estimation for elder financial scams in resource-constrained environments.
Phi-4 and LLaMA-3.2 achieved stronger turn-aware risk estimation performance relative to their parameter scale.
The main limitation is the potential variability in model performance across different scam scenarios and the need for extensive dialogue datasets for broader applicability.
Task-Specific Prompt with Global Context for Multi-Task Graph Pre-Training
Virgil Qiu, Author 2, Author 3 +2 more
TPGC achieves consistently better performance in few-shot settings on 6 benchmarks compared to state-of-the-art baselines, with fewer tunable parameters and lower runtime.
TPGC outperforms state-of-the-art baselines on 6 benchmarks under few-shot settings.
The reliance on an auxiliary graph for prompt initialization may limit generalizability to other graph structures.
From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling
Not specified in the provided content
Circuit-guided weight scaling improves safety rates under adversarial attacks by 26.5% with only a 1.7% drop in accuracy across standard benchmarks.
26.5% improvement in safety rates under attacks across six LLMs.
The paper does not specify the authors or provide detailed experimental setups, which may hinder reproducibility.
Zero-Shot Respiratory Sound Classification through LLM-Augmented Audio-Text Alignment
Not provided
The proposed framework achieves a 61.3% mean zero-shot AUC, outperforming existing models while using only 43% of the data required by full-scale baselines.
Achieved a 61.3% mean zero-shot AUC across 9 tasks on 6 datasets.
The reliance on synthesized structured reports from metadata may limit reproducibility due to potential variability in report quality.
HyperWorld: Hypergraph-Structured State Serialization Improves Learned Textual World Models
Not specified in the provided content
Hyperedge serialization provides the clearest performance gains for learned textual world models, particularly under distribution shift and with limited model capacity.
Hyperedge serialization achieved the highest success rate in downstream greedy planning among tested representations.
The study may have limitations in reproducibility due to the specific model scales and data budgets used in the experiments.
REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent
Not specified in the provided content
REAL-Q reduces end-to-end KL divergence by up to ~49% compared to state-of-the-art globally-guided methods.
Reduces end-to-end KL divergence by up to ~49%.
The paper does not specify authors or provide detailed reproducibility metrics.
Long-Horizon State Tracking in LLMs: Executing MD5 through a Deep Sequence of Dependent Tool Calls
Not provided in the abstract
The model gpt-oss-120b successfully carries the full state across 196 dependent tool calls and returns the correct MD5 digest on a majority of completed runs.
The model maintains state across 196 calls and achieves correct results in a majority of cases.
The study's focus on bookkeeping errors may limit generalizability to other long-horizon tasks.
ReNFT: Repairing Mode Collapse in Reward Post-Training via Internal Probability-Mass Recalibration
Not specified in the provided content
ReNFT successfully repairs a high-reward, low-diversity adapter, improving diversity metrics while retaining nearly all of the original reward.
ReNFT improves DreamSim-Div by 58.8% and 55.0% on PickScore and GenEval, respectively.
The paper does not specify authors or detailed experimental setups, which may hinder reproducibility.
A Cone-Constrained Bilinear Decomposition for Total Scaled-Gradient Variation Models
Not specified in the provided content
The proposed bilinear decomposition method achieves global convergence and improves image restoration performance, particularly under high noise levels.
Achieves PSNR and SSIM competitive with or superior to representative variational methods, especially at high noise levels.
The highly nonconvex and nonlinear nature of the TSGV regularizer may still pose challenges for reproducibility in different contexts.
ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training
Not provided in the abstract
ZimaBlue improves zero-shot evaluation success rates in robotic tasks from 36.1% to 77.8% by leveraging over 120,000 hours of embodied video.
Achieved a 41.7% increase in success rates on real-robot evaluations.
The reliance on large-scale video data may limit reproducibility in environments with less available data.
Discrete-Time MDP Modeling for Multi-Item Capacitated Lot Sizing with Stochastic Demand Timing
Author1, Author2, Author3 +2 more
The proposed genetic algorithm effectively approximates the optimal solution for the stochastic timing problem with an average optimality gap of about 3.44%.
The genetic algorithm achieves an average optimization speedup of 6.89 at the 95% confidence level.
The computational complexity increases significantly with stochastic timing, which may hinder reproducibility on standard hardware.
Behaviorally Grounded User Profiles from the Wild for Personalized Alignment and Multi-Perspective Reasoning
Not provided in the content
Behaviorally grounded profiles significantly improve base models and outperform synthetic profile baselines in personalized language systems.
Profiles consistently improve performance across complex recommendation and open-ended query benchmarks.
The reliance on authentic social media data may limit reproducibility due to data access and privacy concerns.
GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent Environments
Not provided in the abstract
GUI-CC demonstrates that current GUI world models often fail to maintain task-relevant context in multi-step interactions despite producing plausible single-step outputs.
Constructed 500 offline trajectory tasks and 200 emulator-verified online tasks across 30 mobile apps.
The benchmark highlights that plausible single-step generation does not ensure reliable environment simulation, raising concerns about the reproducibility of multi-step interactions.
trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories
Not specified in the provided content
The outcome-only judge detects 84% of loud faults but only 45% of silent faults, indicating a significant blind spot in current evaluation methods.
A step-rubric judge achieves 77% silent recall with zero false alarms at three times the cost of the outcome-only judge.
The study's findings depend on a specific deterministic environment and fault injector, which may limit generalizability.
Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving
Author1, Author2, Author3 +2 more
Qwen-Drive-1.0 achieves strong 3D perception and driving scene understanding while maintaining general vision-language capabilities.
Highly competitive motion-planning performance demonstrated across various evaluation settings.
The staged training recipe may limit reproducibility due to its reliance on specific driving supervision and general-purpose vision-language data.
DISTAL: Distillation and Self-Supervised Pretraining for Structure-Agnostic Materials Property Prediction
Not specified in the provided content
DISTAL achieves the strongest overall performance among all evaluated feature combinations, improving over the reference benchmark on 37 out of 39 benchmark tasks.
The best-performing multimodal configuration combines compositional descriptors, pretrained latent features, and distilled structural features.
The main limitation is the reliance on a pretrained ALIGNN teacher, which may affect reproducibility if the teacher model is not accessible.
Convergence issues in Relational Concept Analysis based on AOC-posets
Author1, Author2, Author3 +2 more
The paper identifies conditions under which convergence can still be ensured in AOC-posets and proposes a convergent variant of the process.
A convergent variant of the process is proposed that preserves the AOC-poset structure.
The proposed method may lead to attributes referring to concepts absent from the final structures, which could affect interpretability.
A Lagrangian View of Flow Matching
Lipman, Liu
The paper establishes a strict invariance condition for optimal single-step generation, leading to a governing quasi-linear advection PDE that explains the straight-line trajectories of Flow Matching.
Solving the governing PDE via the Method of Characteristics yields straight-line trajectories, enabling massive step sizes in generative models.
The requirement for distillation to flatten intersecting characteristics may limit the reproducibility of empirical models.
Distributed Implicit Harm: A Compositional Safety Blind Spot in MLLM-Based Video Moderation
Author1, Author2, Author3 +2 more
The study reveals that existing MLLMs consistently fail to detect Distributed Implicit Harm (DIH) in videos, despite being able to assess individual components correctly.
Developed a dataset of over 9,000 videos to benchmark MLLMs on their ability to detect DIH.
The dataset lacks compositional harm annotations and is difficult to collect due to the nature of DIH.
From GenAI Virtual Patient Dialogue Logs to Teacher-Interpretable Process Evidence: A Learning Analytics Study in Higher Education
Not provided in the abstract
Layered analysis of GenAI VP dialogue logs reveals process patterns associated with high-rated history taking, supporting process-focused feedback in medical education.
Analyzed 1,030 GenAI VP dialogues from 210 medical learners over five weeks.
The educational utility of the coded dialogue logs may vary based on the specific context and implementation in different educational settings.
Memory-Efficient Training-Free Acceleration of Diffusion Transformers with BaryCache
Author1, Author2, Author3 +2 more
The proposed method achieves up to 3.30x end-to-end sampling speedup compared to baseline DiT inference without increasing VRAM footprint.
3.30x end-to-end sampling speedup
The method's performance may vary based on specific implementation details and hardware configurations.
Curvature Cryptanalysis of Smooth Transformer Feed-Forward Networks
Not specified in the provided content
The study establishes that only 16 projected Hessians can recover hidden FFN directions with high alignment, achieving over 93% top-1 agreement in functional extraction.
Achieved an average absolute cosine alignment above 0.94 with only 16 projected Hessians from 8193 black-box queries.
Output rounding and Gaussian noise significantly reduce recovery effectiveness under fixed attack configurations.
A collective capability boundary in frontier large language models on guideline-conformant and case-specific oncology decision-making
Not specified in the provided content
The study found that 42.1% of oncology decision points were answered correctly by none of the evaluated models, highlighting a critical need for architectural changes in LLMs for clinical decision-making.
42.1% of all items were answered correctly by none of the nine models evaluated.
The assumption that any single model can be the sole basis for a clinical decision limits reproducibility and effectiveness.
Statutory AI: Aligning Large Language Models With Legal Norms
Not specified in the provided content
Statutory AI reduced harmful content by 52 to 59 percentage points across tested models, outperforming standard Constitutional AI by approximately 10 percentage points.
Cut computation time by over 50% while reducing harmful content.
The paper does not specify potential limitations or reproducibility concerns.
From Question-First to Analyst-First: Domain-Expert Skills and Verified Knowledge Compilation for Proactive Enterprise Analytics
Not specified in the provided content
The architecture enables a proactive analytics approach by integrating domain-expert skills and an offline knowledge-compilation loop, allowing users to generate insights without needing to formulate questions first.
The system produces durable schema knowledge that drives expert reports, with every published metric re-verified by re-executing its evidence SQL.
The paper does not include user studies or benchmark claims, which may limit the reproducibility of results.
Expert-validated STEM QA
Not specified in the provided content
The creation of a high-quality, expert-validated STEM dataset (N=398) that significantly improved AI model performance by 15% when trained on a private version of the dataset.
Post-training on a private version of the dataset increased performance of the open source model by 15% relative to the baseline model.
The dataset may still have gaps in coverage and potential inaccuracies due to the contest-based data collection and time-bound review process.
MIRAGE-CAD: Construction-Mediated Multimodal Generation of Executable CAD Programs
MIRAGE-CAD achieves 55.4-70.0% build success and 52.3-66.2% STEP export success across different input modalities without retrieval at inference.
Achieved 55.4-70.0% build success on 2,500 held-out queries per modality.
The explicit construction representation may not generalize across all types of geometries and construction methods.
Unsupervised Latent Space Alignment with Hyperspherical Geodesic Matching
Not provided in the content
HGA can effectively align latent spaces in an unsupervised manner, achieving results comparable to supervised methods with minimal or no supervision.
HGA matches supervised results with minimal or no supervision.
The paper does not specify the authors or provide detailed experimental setups, which may hinder reproducibility.
ERR+: Sequential Entropy Resolution for Efficient and Decisive LLM Reasoning
Xrk Arul, Author 2, Author 3 +2 more
ERR+ improves both accuracy and response conciseness in large reasoning models by optimizing the internal reasoning structure through a novel reward mechanism.
Experiments demonstrate consistent improvements across five datasets.
The joint optimization of the two objectives may induce gradient conflict in early training, which could affect reproducibility.
Equivariant Sheaf Neural Networks: Learning Geometric Transport on Graphs
Not provided in the content
ESNN improves dynamics prediction and recovers the gravity axis when symmetry is broken, demonstrating enhanced performance in various applications without requiring higher-order representations.
ESNN yields substantial gains on selected mesh tasks and long-horizon rollouts.
The paper does not specify author names or provide detailed experimental setups, which may hinder reproducibility.
STAGEET: Stage-wise Typed Edit Tagging for Grammatical Error Correction with Arabic as a Case Study
Author1, Author2, Author3 +2 more
STAGEET achieves state-of-the-art results on QALB-2014 while providing a more inspectable correction trajectory through category-aware staged correction.
Achieves state-of-the-art results on QALB-2014.
The framework's complexity may pose challenges for reproducibility and implementation in different contexts.
Understanding Temporal Semantic Stability in Open-Vocabulary UAV Perception through Metric 3D Fusion
Not provided in the abstract
The study reveals that high aggregate world-space agreement can misrepresent temporal stability due to limited repeated-observation support.
The introduction of a voxel-level evaluation framework that characterizes Semantic Belief Drift (SBD), Observation Persistence (OP), and semantic uncertainty.
The findings indicate that conditions reducing world-space recurrence can artificially inflate perceived stability, complicating reproducibility.
NLP-Driven Knowledge Extraction and Thematic Classification of Translated Ancient Indian Medical Texts
Author1, Author2, Author3 +2 more
The study successfully demonstrates how NLP techniques can organize and make accessible the historical medical wisdom of Ayurveda through thematic classification and knowledge graph development.
Thematic classification with BERTopic allows for the identification of underlying medical topics.
The complexity of ancient vocabulary and formats may hinder reproducibility in broader contexts.
Open-Set Cattle Muzzle Identification: A Leakage-Controlled Benchmark and Evaluation Protocol
Not provided in the abstract
The study demonstrates that a hybrid CNN-ViT model can achieve high detection-and-identification rates for cattle muzzle identification while allowing for incremental enrollment.
The hybrid model achieves detection-and-identification rates of 98.3%, 96.4%, and 93.6% at target false-acceptance rates of 10^(-1), 10^(-2), and 10^(-3), respectively.
The performance difference between oracle threshold selection and deployable threshold calibration raises concerns about practical applicability.
DS-Lighting: Making Agent Harnesses Explicit for Data-Science Automation
Author1, Author2, Author3 +2 more
Explicit harness design improves reproducibility, comparability, and reliability in data-science workflows, reducing system-level failures.
Experiments show that explicit harness design reduces avoidable system-level failures in end-to-end data-science workflows.
The reliance on specific open-source benchmarks may limit generalizability across all data-science tasks.
The Halt Vector: Internalizing a Causal Steering Intervention for Efficient Reasoning
Not provided in the abstract
The halt vector intervention reduces reasoning time by about 25% while maintaining accuracy across five unseen benchmarks.
The halt removes about a quarter of the thinking at held accuracy across five unseen benchmarks.
Installing the halt vector intervention is complex and may corrupt off-axis dimensions relied upon by downstream readers.
Parametric Multimodal User Memory: Storing What Captions Cannot Carry
Not provided
The proposed model achieves a recall of 0.96 for identity recognition by combining a vision-language model and a dedicated encoder, outperforming traditional methods.
The combined model reaches a correct-region oracle recall of 0.96.
The model's performance is dependent on the training-free recognition core, which may limit its adaptability to new contexts.
Gurukul AI: An Interactive AI-Driven Educational Platform for Indian Education System
Not specified in the provided content
The authors created a publicly available dataset of 18,720 question-answer pairs aligned with the Indian curriculum and fine-tuned the LLaMA model to develop an interactive educational platform.
The dataset comprises 18,720 question-answer pairs across five subjects.
The paper does not specify the authors, which may limit reproducibility and credibility.
Improving Spatial-Temporal Reasoning in Video-Language Models with Structured Video Prompting
Not provided
Structured video prompting improves performance in video-language models by organizing visual evidence at inference time without altering model weights.
Structured inputs improve performance in several cases across two video benchmarks.
The method is training-free, which may limit understanding of its underlying mechanisms.
Revisiting Local Context for Long-Horizon Streaming 3D Reconstruction
Not provided in the abstract
ABot-Recon achieves a 40% reduction in absolute trajectory error (ATE) and relative pose error (RPE-R) compared to the best prior results on long-sequence benchmarks.
On Oxford Spires, ABot-Recon achieves an ATE of 4.35 m and an RPE-R of $0.12^ heta$.
The abstract does not specify limitations or reproducibility concerns.
Accelerating LLM Inference via Vector Index Based Output Embeddings
Not provided in the content
The proposed method accelerates output projection and improves end-to-end batch-size-one decoding throughput by up to 82% for the Gemma 3 270M model while maintaining generation quality.
Improved decoding throughput by up to 82% for Gemma 3 270M.
The paper does not specify potential limitations or reproducibility concerns.
SciReC: Diagnostic Evaluation of Multimodal, Multi-Turn Relational Reasoning with Adaptive Interaction
Not provided in the content
The study reveals that relational reasoning is the primary source of error across all models, with Claude 4.6 achieving the best overall relational score of 73%.
Claude 4.6 achieved the best performance on the overall relational score with 73%.
Open-source models achieve their lowest scores on spatial relations, indicating potential limitations in their performance across different relational reasoning tasks.
Sledgehammer or Scalpel? A Fine-grained Adaptive Framework for Implicit Hate Speech
Not provided
The Fine-grained Adaptive Implicit Hate speech Detection (FAID) framework significantly outperforms state-of-the-art baselines by focusing computational resources on complex implicit samples.
FAID significantly outperforms SOTA baselines on four benchmark datasets.
The paper does not specify the authors or provide detailed reproducibility metrics.
VidParse: Online Parsing of Egocentric Procedures Like a Pro
Not provided in the content
VidParse achieves up to a 10x improvement in complex multi-step parsing accuracy over strong online baselines without requiring any gradient updates.
10x improvement in complex multi-step parsing accuracy
The framework is training-free, which may limit its adaptability to specific tasks or datasets.
PACE: Publisher-Adaptive Content Extraction via Agentic Automation
PACE outperforms scalable non-manual baselines while achieving extraction quality comparable to manually engineered parsers.
PACE achieves stronger extraction of article text, metadata, images, and tables, demonstrating effective automation of publisher-specific extraction.
The framework's reliance on representative pages and user requirements may limit its applicability across diverse publishers.
Quanta Perception as Probabilistic Events
Not provided in the abstract
The proposed method processes photon streams at over 50,000 quanta frames per second on commodity GPU hardware, achieving outputs up to four orders of magnitude faster than existing methods.
Processes input streams exceeding 50,000 quanta frames per second.
The paper does not specify the authors or provide detailed experimental setups, which may hinder reproducibility.
DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization
Not provided in the abstract
DAMP reduces recurrent-state storage by 69.1% and accelerates the recurrent-state update kernel by up to 2.01x while maintaining average accuracy close to the FP32 baseline.
DAMP achieves 9.9 bits per state value with a 69.1% reduction in storage.
The study focuses on specific language models (GDN and KDA), which may limit generalizability to other architectures.
Marginal Coverage Credit Reduces Redundant Exploration in Parallel State-Entropy Optimization
Not provided in the content
MCC-PGPSE effectively reduces redundant exploration and improves complementary coverage among parallel policies in discrete state spaces.
MCC-PGPSE produced positive final window gains in normalized team state entropy and state support across all tested settings.
The paper does not specify the authors or provide detailed experimental setups, which may affect reproducibility.
A Deeper Analysis of Block-Sparse Featurizers
Fel, Author2, Author3 +2 more
The introduction of a Tournament Top-K selection rule significantly reduces feature splitting in block-sparse featurizers.
The proposed changes lead to improved performance metrics, specifically reducing feature splitting by a notable percentage.
The BSF still suffers from classic SAE failure modes, which may affect reproducibility.
FVeinSyn: Synthetic Finger Vein Image Generator
Evan Wang, Author 2, Author 3 +2 more
FVeinSyn generated 500,000 synthetic finger vein images, significantly improving model accuracy by 27.43% compared to real-data-only baselines.
Generated 500,000 images with 10,000 vein identities and 50 samples per identity.
The main limitation is the reliance on synthetic data, which may not fully capture the complexities of real-world scenarios.
Quantization-Triggered Backdoors in Language Models: Cross-Quantizer Transferability and the Validation--Deployment Gap
Not provided in the abstract
The study proves that quantization can trigger backdoor attacks in language models, demonstrating that source-precision certification does not ensure behavioral equivalence in deployed configurations.
Backdoored translation models showed up to 85.02% inversion after quantization.
The study's findings may vary across different quantization schemes and model architectures, complicating reproducibility.
The Effect of Emotional Context on Large Language Models' Endorsement of Premature Decisions: Comparing Emotional Vulnerability Across Six Commercial Models
Author1, Author2, Author3 +2 more
Emotional context significantly increases the endorsement of premature decisions by large language models, with five out of six models showing a notable effect.
Emotional expression increased endorsement strength from 18.6 (neutral) to 31.5 (distress), a difference of +12.9 points.
The study's findings may not generalize across all models or emotional contexts, as only six models were tested.
Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning
Not provided in the content
Code-as-World provides a scalable foundation for physical intelligence by achieving state-of-the-art performance on quantitative physical reasoning tasks.
Code-as-World-VL surpasses leading proprietary models on the QuantiPhy benchmark.
The paper does not specify the authors or provide detailed experimental setups, which may hinder reproducibility.
When Muon Meets Task Interference: A Spectral Perspective on Continual Learning and Model Merging
Author1, Author2, Author3 +2 more
The Muon optimizer improves accuracy by up to +5.02 points on model-merging benchmarks and delivers consistent gains in continual learning tasks.
Replacing the AdamW optimizer with Muon improves accuracy by up to +5.02 points.
The paper does not address potential limitations in the generalizability of the Muon optimizer across all tasks.
SLM-Conditioned Hierarchical Relation Routing for Labeled Property Graph Learning
Not provided in the abstract
The proposed architecture allows for dynamic message routing based on semantic evidence, improving prediction accuracy in labeled property graphs.
The architecture supports interpretable analysis at both the neighbor and relationship-type levels.
The integration of a small language model may introduce complexity that affects reproducibility.
TreeGraft: Adaptive Multi-Drafter Grafting for Tree-Based Speculative Decoding
Author1, Author2, Author3 +2 more
TreeGraft outperforms fixed single-drafter strategies by an average of 15.1%, with a maximum gain of 26.6% across various benchmarks.
Achieved a 15.1% average improvement over single-drafter strategies.
The reliance on a lightweight scheduler may limit generalizability across different model architectures.
A Unified Framework for the Mechanics of Information in Convolutional Neural Network Image Space
Not provided in the content
The paper establishes a correspondence between discrete filter symmetry in CNNs and the relativistic energy-momentum relation, enhancing the understanding of information propagation in image processing.
Demonstrations in 3D images reveal blob-like, scale-invariant Morse critical points across various physical scales.
The paper does not specify practical implementations or reproducibility details for the proposed framework.
Surgical Video Generation From Diffusion to World Models: A Survey
Surgical video generation can significantly enhance training and simulation by providing a structured approach to data scarcity in surgical contexts.
The survey categorizes methods into unconditional generation, conditional generation, and world modeling generation, revealing a shift towards modeling causal dynamics.
There is a persistent gap between pixel-level fidelity and clinical plausibility, which poses challenges for reproducibility.
Can a Model Catch Its Own Hallucinations for Free?: Label-Free Doubt Signals Hold Their Own Against a Labelled Dataset for Abstention
Author1, Author2, Author3 +2 more
The study demonstrates that a model's own confidence can effectively replace labelled datasets for training abstention in factual question answering.
The label-free method matches the performance of label-supervised abstention-tuning across six open-weights models.
The method struggles with confidently wrong facts, which it cannot identify as uncertain.