arxiv:2609.21996
Hiskias Dingeto
hisku
AI & ML interests
NLP, Meta-Learning, AI Safety, Mechanical Interpretability
Recent Activity
authored a paper 1 day ago
A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal authored a paper 1 day ago
Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations submitted a paper 2 days ago
A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal