TAPe+ML: A Compact Structured Representation for Multi-Task Computer Vision
Abstract
We present TAPe+ML v3, a compact computer vision system based on TAPe (Theory of Active Perception), a structured representation that encodes relations among perceptual elements before recognition. Instead of operating directly on pixel tensors, the system uses a shared TAPe representation and a modular recognition architecture for image classification, object detection, and instance segmentation. TAPe+ML v3 combines background and contour processing, local object localization, prototype-based classification, and a coordinator for specialized submodels. Across the reported experiments, it uses fewer than 100,000 parameters. On COCO object detection, it obtains 84.7 mAP50 and 65.3 mAP50-95. On COCO instance segmentation, it obtains 80.7 mask mAP50 and 58.4 mask mAP50-95. In classification experiments, it reaches 92 percent validation accuracy on Imagenette under an identical-training comparison with a raw-pixel baseline, and 89.9 percent Top-1 accuracy on ImageNet-Real. We also evaluate compactness in video scene detection and adaptation under distribution shift in an industrial pilot. The results suggest that shifting part of the modeling burden from network parameters to a structured input representation can support compact multi-task vision systems with reduced data, memory, and compute requirements.
Community
TAPe+ML v3 is a compact computer vision architecture for classification, object detection, and instance segmentation.
The paper reports 88.1% ImageNet-1k Top-1 accuracy, 84.7 mAP50 and 65.3 mAP50-95 on COCO detection, and 80.7 Mask mAP50 and 58.4 Mask mAP50-95 on COCO instance segmentation, with fewer than 100K parameters.
It also includes experiments on TAPe data versus raw pixels, video scene detection, auto-annotation, and an industrial domain-shift pilot.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- HiPerViT: A Hierarchical Perceiver-Vision Transformer Architecture for Multi-Scale Texture Recognition (2026)
- A Controlled Study of Self-Supervised Image and Video Pretraining under Limited Resources (2026)
- When Simplicity Wins: Bottleneck-Aware Context Modeling for Lightweight Semantic Segmentation (2026)
- Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding (2026)
- Hardware-Aware Learned Representation Compression for Distributed In-Sensor Vision (2026)
- WALDO: One-Shot Exemplar-Conditioned Object Detection in Cluttered Scenes (2026)
- FAVE: Foveated Adaptive Visual Encoding for Efficient Fine-Grained Visual Understanding (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper