RenikudPlus — Hebrew grapheme-to-phoneme
Hebrew text in, IPA out, with stress marks — the part unpointed Hebrew leaves to context. For TTS front-ends, lexicon work and pronunciation research.
Results
Target-word accuracy on 1,653 sentences of held-out Hebrew:
| RenikudPlus | Gemini | ReNikud (paper) | Phonikud | |
|---|---|---|---|---|
| OVERALL | 88.6% | 84.6% | 78.2% | 66.4% |
| Gender | 99.3% | 77.0% | 59.9% | 37.5% |
| Min. Stress Pairs | 93.3% | 88.0% | 80.7% | 78.7% |
| Stress Homographs | 92.6% | 92.6% | 79.8% | 78.3% |
| Names | 80.0% | 75.3% | 67.3% | 68.0% |
| Acronyms | 82.2% | 79.0% | 56.6% | 37.5% |
| Slang | 82.7% | 71.2% | 59.0% | 41.7% |
| Penultimate Stress | 88.1% | 89.4% | 82.1% | 60.9% |
| Rare Phonemes | 76.8% | 57.6% | 41.1% | 19.9% |
| Foreign | 87.1% | 71.6% | 56.8% | 36.1% |
| Colloquial | 51.0% | 25.2% | 54.3% | 9.3% |
| ILSpeech-test | 93.7% | 96.0% | 92.5% | 85.5% |
Usage
pip install onnxruntime numpy
Download renikud_onnx.py and a model file from this repo into the same folder:
from renikud_onnx import G2P
g2p = G2P("model_int8.onnx")
g2p.phonemize("שלום, מה נשמע?")
# ʃalˈom, mˈa niʃmˈa?
g2p.phonemize("הלכתי לספר וקראתי ספר בזמן שהוא ספר כמה אצבעות יש לו")
# halˈaχti lasapˈaʁ vekaʁˈati sˈefeʁ bizmˈan ʃehˈu safˈaʁ kˈama ʔetsbaʔˈot jˈeʃ lˈo
# barber book counted
Four occurrences of ספר, four different words, each resolved from context.
Speaker and addressee
Hebrew inflects for who is speaking and who is addressed; the text usually shows neither.
Both controls default to 0 (unknown), 1 = male, 2 = female.
g2p.phonemize("אני רוצה להגיד לך משהו חשוב", speaker=2, target_speaker=1)
| speaking → addressing | אני רוצה להגיד לך משהו חשוב |
|---|---|
| man → man | ʔanˈi ʁotsˈe lehaɡˈid leχˈa mˈaʃu χaʃˈuv |
| man → woman | ʔanˈi ʁotsˈe lehaɡˈid lˈaχ mˈaʃu χaʃˈuv |
| woman → man | ʔanˈi ʁotsˈa lehaɡˈid leχˈa mˈaʃu χaʃˈuv |
| woman → woman | ʔanˈi ʁotsˈa lehaɡˈid lˈaχ mˈaʃu χaʃˈuv |
speaker sets רוצה, target_speaker sets לך. Four readings of one unchanged sentence.
Numbers and digits
Digits are expanded to Hebrew words before phonemization, with gender agreement taken from the counted noun:
g2p.phonemize("יש לי 3 ילדים") # jˈeʃ lˈi ʃloʃˈa jeladˈim
g2p.phonemize("יש לי 3 בנות") # jˈeʃ lˈi ʃalˈoʃ banˈot
g2p.phonemize("המחיר 1250 שקלים") # hameχˈiʁ ʔˈelef matˈajim veχamiʃˈim ʃkalˈim
Files
| file | what |
|---|---|
model.onnx |
fp32 |
model_int8.onnx |
quantized, 4× smaller, faster on CPU |
renikud_onnx.py |
wrapper — decoding, long-input windowing, number expansion |
Notes
- Stress is a first-class output (
ˈbefore the stressed syllable). - Remaining errors concentrate in unattested foreign names and in transcription-convention choices
(glottal-stop realisation,
ba/beproclitics) where more than one reading is correct Hebrew.
