Evaluating the Effects of Prompt Perturbation on Bias and Hallucination in Large Language Models
Author1, Author2, Author3, Author4, Author5
Abstract
This study evaluates the robustness of large language models to prompt perturbations, revealing that such variations can mitigate bias and hallucination in some models, with Claude 3 outperforming others like GPT3.5.
Reality Card
Perturbed variations of prompts can reduce bias and hallucination in large language models, with Claude 3 showing superior performance in decision-making tasks compared to GPT3.5.
Claude 3 is more effective for the tasks represented in most datasets, outperforming GPT3.5 in several cases.
The study highlights the need for rigorous testing and validation, indicating potential variability in model performance across different contexts.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.