AI Papers This Week
Top 10 arXiv papers from the past 7 days, ranked by builder relevance. Core claim, method highlight, and limitations — distilled into 30-second reads.
Synergising Local Geo-Environmental Characteristics with Spatial Context for Enhancing Landslide Susceptibility Mapping
Not specified in the provided content
The LGSCF-based models significantly improve landslide susceptibility mapping accuracy, achieving F1-scores up to 87.09% and AUC values up to 0.9472.
Achieved F1-scores up to 87.09% and AUC values up to 0.9472.
The study does not specify potential limitations or reproducibility concerns.
SHIFT-LLM: Distribution Shift Correction in Depth-Pruned LLMs
Not specified in the provided content
SHIFT-LLM recovers accuracy lost to depth pruning by using Linear Residual Adapters, achieving gains up to +15.7 points on Llama-3.1-8B-Instruct.
Achieves accuracy recovery with only a few hundred calibration samples and no gradient computation.
The method relies on a small held-out set for calibration, which may limit generalizability across different datasets.
Fusing Perceptual Vision Experts with Multimodal Large Language Models for Explainable Plant Disease Diagnosis: From Benchmark Imagery to Real-World Robotic Field Validation
Not specified in the provided content
The Hybrid Hierarchical Multi-Agent Framework (H$^{2}$MAF) improves plant disease diagnosis accuracy significantly, achieving up to 99.3% accuracy in real-world conditions.
Gemma improves accuracy from 63.9% to 68.5% on the PlantDoc dataset.
Calibration dependency of MLLM arbitration may affect reproducibility.
A Lightweight Multimodal Vision-Language Framework for Early-Stage Anatomical Green Fruit Classification in Commercial Orchards
William Bu, Author 2, Author 3 +2 more
The framework achieved high F1-scores (0.95 for calyx, 0.98 for fruitlet, and 0.85 for peduncle) using a lightweight model optimized for edge deployment.
Achieved a macro-F1 score of 0.93 with millisecond-level patch inference on NVIDIA Jetson hardware.
The dataset is limited to 600 images from specific orchards, which may affect generalizability.
Dynamic Influence-Weighted Distillation for Single-IMU Activity Recognition
Not provided in the abstract
Dynamic Influence Weighting (DIW) improves the performance of a single-IMU model by achieving a pooled out-of-fold macro-F1 score of 0.638451, surpassing both Supervised and Fixed-weight KD methods.
DIW achieved a 7.66 percentage point improvement over Supervised methods in activity recognition.
The study relies on a specific dataset (WEAR) and may not generalize to other datasets or real-world applications without further validation.
Targeting the Attention Heads Behind Object Hallucination in LLaVA
Author1, Author2, Author3 +2 more
The proposed method reduces the fraction of captions with hallucinated objects from 0.370 to 0.230 and hallucinated object mentions from 0.156 to 0.096.
The combined method lowers hallucination rates significantly on 400 held-out COCO images.
The method may lower object recall from 0.78 to 0.70, indicating a trade-off between hallucination reduction and recall.
The Dialect Tax: Dialectal Biases Persist throughout the Language Modeling Pipeline
Author1, Author2, Author3 +2 more
The study reveals that dialectal biases are encoded and accumulated at every step of the language modeling process, affecting model performance and accuracy.
During pre-training, dialect pairs induce more divergent gradient updates compared to unrelated SAE documents.
The findings indicate that dialectal performance gaps are complex and persistent, making reproducibility challenging across different model families.
ExFold: Unified Expert Folding for Training-Free MoE Prefill-Decode Acceleration
Author1, Author2, Author3 +2 more
ExFold achieves up to 1.41x TTFT and 2.45x TPOT speedups while retaining about 99% of the original average quality.
1.41x TTFT and 2.45x TPOT speedups.
The method relies on calibrated scalar projectors which may require careful tuning on different datasets.
GreenLeaf Law Embed Tiny: A Compact Embedding Model for Legal Domain Retrieval
GreenLeaf-Tiny achieves 75.11% on the Massive Legal Embedding Benchmark (MLEB) while being a 0.6B parameter model.
Utilizes a two-stage training pipeline and a dataset of 3.4 million query-passage pairs.
The model's performance may be limited by the quality and diversity of the training dataset.
EduRiskX: A Neuro-Symbolic Framework with F-Logic Reasoning for Early Academic Risk Prediction
Not provided
EduRiskX achieves an accuracy of 0.900 and an F1-score of 0.894 for early academic risk prediction, with a detection rate of 94.30 percent.
Achieved an average early detection week of 9.32.
The paper does not specify the authors, which may hinder reproducibility.
Get the weekly paper digest in your inbox
Every Monday, the top 10 arXiv papers ranked by builder relevance — with core claim, method, and limitations. No fluff. Just the signal.
The Signal Brief
The only AI brief that separates confirmed facts from official claims — and tells you what actually changed.
Role-aware. No scroll trap. Every morning.