Abstract
The Conformal Relevance framework automates score function design for conformal prediction via in-context learning and ensembling to improve conciseness while preserving coverage across NLP retrieval tasks.
Many NLP tasks, such as summarization and extractive question answering, reduce to retrieving relevant content from documents under two constraints: coverage, retaining enough pertinent information to achieve some goal, and conciseness, removing as much irrelevant information as possible. Conformal prediction methods have been used to guarantee coverage, and must be optimized for conciseness through design of a score function. State-of-the-art scoring functions use hand-engineered LLM prompts asking the model to rate the importance of content, but manual prompt engineering is labor-intensive and task-specific. We introduce the Conformal Relevance framework which uses in-context learning example curation and ensembling to create a score function which maintains coverage while improving conciseness with minimal manual input. We demonstrate this framework's application on seven NLP tasks, and also theoretically study the impact of diversity for ensembled conformal scores, giving a complementarity condition that characterizes when ensembling improves worst-case sentence scores, and a saturation bound on ensemble improvement.
Community
Previously NLP tasks were tackled one-by-one with specially crafted LLM prompts to create scoring functions. We show that many tasks can be conformalized with a single prompt format based on ICL, rather than manual crafting.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- EdgeLM: Edge Demonstrations for Language Models'Table Understanding (2026)
- Conformalized Large Language Models under Configuration Shift (2026)
- Decoupling Generation and Selection for Budget-Constrained Faithful Summarization (2026)
- Prompt-Robust Language Models: Which Training Strategies Work? (2026)
- From Preferences to Principles: Rubric-Based Alignment for Grounded Knowledge Answers (2026)
- Hierarchical Group-Conditional Conformal Risk Control for Selective Prediction in Language Models (2026)
- QUBO-Optimized Evidence Selection for Retrieval-Augmented Question Answering with Unconventional Solvers (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.03005 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper