evidenceprune-modernbert-large

evidenceprune is a lightweight context-pruning model fine-tuned for fact checking. Given a claim and a web page, it marks the sentences a fact-checker would need.

This checkpoint is a 0.4B encoder fine-tuned from ModernBERT-large as a token classifier.

Usage

pip install git+https://github.com/ofbread/evidenceprune
from evidenceprune import Pruner

kept = Pruner("ofbread/evidenceprune-modernbert-large").prune(claim, documents)

Input

Pruner(model, threshold=None, device=None) loads the model. prune(claim, documents, requirements=None, claim_date="", speaker="", cap_windows=16, stop_after_empty=2, doc_ceiling=12000) reads the pages.

parameter type meaning
claim str the claim being checked
documents list of {"text": str, "title": str, "url": str} the pages; title and url are optional but the model was trained with them
claim_date str, optional when the claim was made, YYYY-MM-DD
speaker str, optional who made the claim
requirements list of str, optional what must be established to check the claim
threshold float, default 0.3458 a sentence is kept when its P(keep) is at or above this
cap_windows int, default 16 read at most this many windows of 48 sentences (at most 10,000 characters each) per page
stop_after_empty int, default 2 stop reading a page after this many consecutive windows with nothing kept
doc_ceiling int, default 12000 a page contributes at most this many characters
device str, optional cuda or cpu

Output

prune returns one result per document, in order:

field type meaning
text str the kept sentences, in page order, gaps marked with […]
kept list of (start, end) offsets of each kept sentence in the document's text
sentences list of str the same sentences as strings
probs dict, sentence index → float P(keep) for every sentence the model read
title_kept bool the title is read as the page's first sentence
windows_read, n_windows int windows the model read, windows in the page
stopped_early bool whether the stop rule ended the read
Downloads last month
7
Safetensors
Model size
0.4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ofbread/evidenceprune-modernbert-large

Finetuned
(370)
this model