Papers/2608.24936
🧪 Test?View on arXiv

GreenLeaf Law Embed Tiny: A Compact Embedding Model for Legal Domain Retrieval

embeddinglegal retrievalfine-tuningknowledge distillation
2608.24936
Builder Relevance
80%
3h ago

Abstract

A compact embedding model for legal domain retrieval that achieves competitive performance with limited parameters.

Reality Card

Core Claim

GreenLeaf-Tiny achieves 75.11% on the Massive Legal Embedding Benchmark (MLEB) while being a 0.6B parameter model.

Method / Result

Utilizes a two-stage training pipeline and a dataset of 3.4 million query-passage pairs.

Limitations

The model's performance may be limited by the quality and diversity of the training dataset.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers