nm-testing/llama2.c-stories15M-ultrachat-mixed-uncompressed 24.4M • Updated about 1 month ago • 2.16k
nm-testing/llama2.c-stories42M-gsm8k-quantized-only-uncompressed 58.1M • Updated about 1 month ago • 2.49k
nm-testing/TinyLlama-1.1B-Chat-v1.0-W8A8-Dynamic-Per-Token-uncompressed 1B • Updated about 1 month ago • 400
nm-testing/llama2.c-stories15M-pruned_50.2of4-uncompressed-tensor_weights_tensor_act_fp8-BitMaskCompressed 22.2M • Updated about 1 month ago • 27
nm-testing/SparseLlama-3.1-8B-gsm8k-pruned.2of4-quantized.w4a16-uncompressed 2B • Updated about 1 month ago • 18
nm-testing/TinyLlama-1.1B-Chat-v1.0-pruned_50.2of4-FP8-uncompressed 1B • Updated about 1 month ago • 16
nm-testing/tinyllama-oneshot-w8w8-test-static-shape-change Text Generation • 1B • Updated about 1 month ago • 6.72k
nm-testing/tinyllama-oneshot-w8a8-dynamic-token-v2-asym Text Generation • 1B • Updated about 1 month ago • 173
nm-testing/tinyllama-oneshot-w8a8-dynamic-token-v2 Text Generation • 1B • Updated about 1 month ago • 2.42k
nm-testing/tinyllama-oneshot-w8a8-channel-dynamic-token-v2 Text Generation • 1B • Updated about 1 month ago • 2.65k
nm-testing/tinyllama-oneshot-w8a16-per-channel Text Generation • 1B • Updated about 1 month ago • 1.3k