Integrating Fairness and Explainability in a Multiple Instance Reinforcement Learning System
Not provided in the abstract
Abstract
This study investigates a multi-objective framework that combines reinforcement learning-based multiple instance learning (RL-MIL), adversarial debiasing, and preference-conditioned hypernetworks for student-at-risk prediction.
Reality Card
The study demonstrates that fairness objectives can be integrated into an interpretable RL-MIL pipeline, but preference conditioning does not ensure controllable multi-objective behavior.
The underlying RL-MIL baseline achieves strong classification performance, but hypernetwork extensions exhibit mode collapse.
The main limitation is the mode collapse in hypernetwork extensions, which affects the systematic movement along the fairness-performance frontier.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.