T-Students FITDNU is a dataset for object/behavior detection in classroom environments, collected using a single ceiling-mounted camera (1080p@25fps) during real class sessions (morning/afternoon). The dataset emphasizes small objects (phones, computers) and occlusions due to crowded classrooms, making it valuable for evaluating YOLO models in real-world deployment scenarios.
- 9 classes:
using_phone,using_computer,sleeping,turning_left,turning_right,raising_hand,writing,phone,computer. - Scale: 3,351 images, 126,429 bounding boxes.
- Goal: Serve research and real-time demo purposes from a single ceiling viewpoint.
Bounding box distribution by class:
| No. | Class | # of bbox |
|---|---|---|
| 1 | Computer | 20,100 |
| 2 | Phone | 45,910 |
| 3 | Raising Hand | 318 |
| 4 | Sleeping | 4,826 |
| 5 | Turning Left | 9,720 |
| 6 | Turning Right | 9,222 |
| 7 | Using Computer | 8,412 |
| 8 | Using Phone | 23,709 |
| 9 | Writing | 4,212 |
The dataset is intentionally imbalanced to reflect natural frequency — useful for Focal Loss, re-weighting, oversampling, copy-paste, etc.
Roboflow Universe: open the dataset page (Link), choose a Version and Export Format (YOLOv5/YOLOv8/COCO JSON/VOC…), and download via UI or API.
- YOLOv8l achieves the highest overall mAP@[0.5:0.95], especially excelling in subtle behaviors or small objects (sleeping, turning_left/right, using_computer).
- YOLOv7 performs well on frequent classes (using_phone), but performance drops on rare ones (raising_hand).
- YOLOv12s is a lightweight and fast option, ideal for real-time deployment.
- Faster R-CNN struggles with occlusion and crowded classroom scenes, leading to low mAP in phone and raising_hand classes.
- Dataset/Paper: [email protected]
- For technical issues: open an Issue in this repository.
