Tiny-Jev support research preview: rl
Independent project. Not affiliated with or endorsed by TypeSafe AI (makers of Jev) or the CLM authors. "Jev" is used only to describe the interface style this research studies.
Experimental research preview maintained by sainath.
Repository: sainath/tiny-jev-support-v03-rl. Code and trained decision-head weights: Apache-2.0.
Independent human evaluation has not been performed.
What is included
1,123,073 decision-head parameters, trained on synthetic support data with frozen
Qwen3-1.7B representations. The 1.7B backbone is required and downloaded separately
at revision 70d244cc86ccca08cf5af4e1e306ecf908b1ad5e. This is not a standalone 1.1M model.
Uses a custom Python loader, not AutoModelForCausalLM or an automatic hosted endpoint.
Inspired by Jev's typed interface; no affiliation with TypeSafe or CLM.
Method and provenance
Variant: reinforce. Seed: 42. Selected epoch: 9.
Fixed seed 42 for a compact preview, retaining both comparison methods; not selected by test ranking.
The supervised control directly optimizes Brier loss; the RL comparator uses
REINFORCE with Gaussian latent reports, eight samples and a leave-one-out baseline,
rewarded by negative multiclass Brier. Both start from the same supervised initialization.
This does not reproduce Jev's RLCD. RL superiority is not supported by the results.
Checkpoint and weight hashes, encoder settings and data hash are in config.json.
Trainable heads were selected on validation; calibration/test rows were excluded from training.
Evaluation for these exact weights
Dataset: support-varied-v2 (combined wording/order variation), 2,000 synthetic scenarios / 8,000 questions. Held-out test: 200 scenarios / 800 questions. AI-authored paraphrases; no independent human review. Serving temperature: 1.0.
| Metric | Uncalibrated test result |
|---|---|
| Accuracy | 92.50% |
| Multiclass Brier (lower is better) | 0.105198 |
| NLL (lower is better) | 0.158059 |
evaluation.json also contains calibration ablations; those settings are not automatically enabled in serving. three_seed_results.json describes the wider experiment, not just this checkpoint. These test results have informed development and are exploratory. Do not compare different test distributions as a causal gain.
Known limits
Generalization to independently written text and unseen tasks is unestablished. BANKING77 evaluation of earlier, different weights was near chance; these new weights have not yet been evaluated on BANKING77. Calibration on synthetic data does not establish real-world calibration. Escalation is a simulated outcome task. Choice/Noul/Score outputs are typed; this does not guarantee correct answers. Questions are independently batched and the state is encoded for each question. Inputs exceeding the saved encoder limit (256 tokens including formatting) are rejected.
Installation and use
Download this repository, then in a GPU Python environment run:
python -m pip install ./tiny_jev-0.4.0-py3-none-any.whl
from tiny_jev.inference import TinyJev
from tiny_jev.schema import Question
model = TinyJev("/path/to/downloaded/repository", device="cuda:0")
result = model.decide("I explicitly request a refund.", {
"refund": Question.noul("Does the customer explicitly request a refund?")
})
print(result)
After installing the included wheel, Hub loading is also supported:
TinyJev.from_pretrained("sainath/tiny-jev-support-v03-rl", revision="FULL_COMMIT_SHA", device="cuda:0").
Replace FULL_COMMIT_SHA with the full revision from this repository's commit history.
The loader downloads data-only config and weights; the Python package is installed explicitly.
Verification and contributions
Local safetensors export is exact. On 2026-09-29 this repository was downloaded
from the Hugging Face Hub (commit 04e18361fd25f67ab153bf60611994d549d408c8) into a clean Kaggle session and
verified end to end with the real Qwen backbone on Tesla T4 (Python 3.12.13,
PyTorch 2.10.0+cu128, tiny-jev 0.4.0). Weight hash and encoder metadata match
this package; see verification.json.
Eight synthetic questions passed the serving smoke/regression check, including typed probability validity and batched/single-question agreement. Maximum absolute probability difference from saved predictions: 0.000411890 (allowed: 0.01). This is not a new accuracy benchmark or human evaluation. Both the native serving path and the v0.4 Jev-style adapter passed. See FEEDBACK.md to submit voluntary reproducible reports.
Jev-style application testing (library v0.4)
This package adds a documented request subset for comparative application testing. It is not a drop-in TypeSafe SDK or HTTP API replacement. See JEV_COMPATIBILITY.md for supported inputs and limitations.
from tiny_jev.compat import JevAdapter
client = JevAdapter.from_release("/path/to/downloaded/repository", device="cuda:0")
response = client.system_one(
state="I explicitly request a refund.",
questions={"refund": {"type": "noul", "instructions": "Does the customer explicitly request a refund?"}},
)
print(response["answers"]["refund"]["noul"])
No confidence or billed usage fields are fabricated. Existing applications
depending on those fields need explicit changes. Choice/Score diagnostics use
maximum probability and entropy, not an asserted copy of Jev confidence.
The comparison runner (python -m tiny_jev.compare) evaluates local models or
saved Jev responses offline; live paid Jev calls require explicit --call-jev.
Example fixtures are existing synthetic cases, not independent benchmark evidence.
The adapter has contract tests, a CPU integration check using actual weights and
cached Qwen vectors (adapter_verification.json), and passed the end-to-end
real-Qwen check on the hosted copy (verification.json, adapter_status: passed).
Attribution
Backbone: https://huggingface.co/Qwen/Qwen3-1.7B (Apache-2.0). Interface inspiration: https://typesafe.ai/blog/introducing-system-one-models-and-jev Frozen-backbone inspiration: https://github.com/Contrastive-LM/CLM This implementation and its documentation were developed with substantial AI assistance. Source and tests are included in source.zip. The original project code, documentation, synthetic verification examples and trained heads are licensed under Apache-2.0; see LICENSE and NOTICE. Third-party dependencies and separately downloaded datasets retain their own licenses. Qwen weights are downloaded separately and are not included in this repository. The wheel includes the same license and notices. Version 0.4 adds the comparison adapter; the backbone recipe, native inference code and decision-head weights are unchanged.
- Downloads last month
- 12