SpatialBlock-4B-direct-GGUF

SpatialBlock-4B-direct is an open-source multimodal model released on Hugging Face by rsoohyun under the Apache-2.0 license, developed to enhance spatial intelligence in Large Vision-Language Models (LVLMs). Based on the Qwen/Qwen3-VL-4B-Instruct base architecture and supported by the Hugging Face transformers library via the image-text-to-text pipeline, this checkpoint is fine-tuned on the synthetic SpatialBlock-15k dataset as presented in the paper SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem. It specializes in directly predicting solutions to intricate spatial tasks—including 3D-to-2D projection, viewpoint transformation, and structural combination—with complete training methodology, evaluation metrics, and the companion “reason” model accessible via its GitHub repository.

Model Files

File Name Quant Type File Size File Link
SpatialBlock-4B-direct.BF16.gguf BF16 8.05 GB Download
SpatialBlock-4B-direct.Q4_K_M.gguf Q4_K_M 2.5 GB Download
SpatialBlock-4B-direct.Q5_K_M.gguf Q5_K_M 2.89 GB Download
SpatialBlock-4B-direct.mmproj-bf16.gguf mmproj-bf16 839 MB Download

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Downloads last month
174
GGUF
Model size
4B params
Architecture
qwen3vl
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/SpatialBlock-4B-direct-GGUF

Quantized
(1)
this model

Collection including prithivMLmods/SpatialBlock-4B-direct-GGUF