Commit ·
bef89c9
1
Parent(s): 402ede5
Upload 3 files (#1)
Browse files- Upload 3 files (e392384e5bc4761815dd34073a32fd1293dfe579)
- README.md +202 -131
- config.json +37 -0
- mc3-18-hmdb51-ucf-transfer.pth +2 -2
README.md
CHANGED
|
@@ -1,4 +1,6 @@
|
|
| 1 |
---
|
|
|
|
|
|
|
| 2 |
tags:
|
| 3 |
- video-classification
|
| 4 |
- action-recognition
|
|
@@ -9,178 +11,229 @@ tags:
|
|
| 9 |
- pytorch
|
| 10 |
- computer-vision
|
| 11 |
- spatiotemporal
|
| 12 |
-
-
|
| 13 |
-
|
|
|
|
| 14 |
datasets:
|
| 15 |
- hmdb51
|
| 16 |
- ucf101
|
| 17 |
metrics:
|
| 18 |
- accuracy
|
| 19 |
-
- f1
|
| 20 |
- precision
|
| 21 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 22 |
---
|
| 23 |
|
| 24 |
-
|
| 25 |
-
|
| 26 |
-
|
| 27 |
|
| 28 |
-
MC3-18
|
| 29 |
|
| 30 |
-
|
| 31 |
|
| 32 |
-
|
| 33 |
|
| 34 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 35 |
|
| 36 |
-
-
|
| 37 |
-
|
| 38 |
-
|
| 39 |
-
|
| 40 |
-
|
| 41 |
-
|
| 42 |
-
- **Output:** 51-class action predictions
|
| 43 |
|
| 44 |
-
|
| 45 |
-
|
| 46 |
-
|
| 47 |
-
|
| 48 |
-
|
| 49 |
-
Batch Size: 16
|
| 50 |
-
Epochs: 100
|
| 51 |
-
Learning Rate: 0.0003
|
| 52 |
-
Weight Decay: 2e-3
|
| 53 |
-
Optimizer: SGD (momentum=0.9)
|
| 54 |
-
```
|
| 55 |
|
| 56 |
-
|
| 57 |
-
- MixUp (alpha=0.4)
|
| 58 |
-
- CutMix (alpha=0.8)
|
| 59 |
-
- Label Smoothing (0.1)
|
| 60 |
-
- Random horizontal flip
|
| 61 |
-
- Color jitter
|
| 62 |
-
- Random grayscale
|
| 63 |
|
| 64 |
## Performance
|
| 65 |
|
| 66 |
-
| Metric
|
| 67 |
-
|--------|-------|
|
| 68 |
-
|
|
| 69 |
-
|
|
| 70 |
-
|
|
| 71 |
-
|
|
| 72 |
-
|
|
| 73 |
-
|
| 74 |
-
## Why UCF-101 Transfer?
|
| 75 |
|
| 76 |
-
UCF-101
|
| 77 |
-
- Both are human action recognition datasets
|
| 78 |
-
- Similar video sources (YouTube, movies)
|
| 79 |
-
- Overlapping action categories (basketball, biking, diving, etc.)
|
| 80 |
-
- Similar temporal and spatial scales
|
| 81 |
|
| 82 |
-
|
| 83 |
-
|
| 84 |
-
## Better Generalization
|
| 85 |
|
| 86 |
-
|
| 87 |
|
| 88 |
-
|
| 89 |
-
|---------------|---------|-----------|-----|
|
| 90 |
-
| Kinetics-400 | 56.34% | ~75% | 19% |
|
| 91 |
-
| UCF-101 (this) | 55.46% | 68.59% | 13% |
|
| 92 |
|
| 93 |
-
|
| 94 |
-
- Similar validation accuracy (only 0.88% lower)
|
| 95 |
-
- Much better generalization (6% smaller train-val gap)
|
| 96 |
-
- Lower training accuracy (less memorization)
|
| 97 |
|
| 98 |
-
|
| 99 |
|
| 100 |
-
##
|
| 101 |
|
| 102 |
-
|
| 103 |
|
| 104 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 105 |
|
| 106 |
-
|
| 107 |
|
| 108 |
-
## Usage
|
| 109 |
```python
|
|
|
|
| 110 |
import torch
|
| 111 |
-
from
|
| 112 |
-
from
|
| 113 |
-
|
| 114 |
-
|
| 115 |
-
|
| 116 |
-
|
| 117 |
-
|
| 118 |
-
|
| 119 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 120 |
model.eval()
|
|
|
|
| 121 |
|
| 122 |
-
#
|
| 123 |
-
|
| 124 |
-
|
| 125 |
-
|
| 126 |
-
|
| 127 |
-
|
| 128 |
-
|
| 129 |
-
|
| 130 |
-
|
| 131 |
-
|
| 132 |
-
|
| 133 |
-
frames = [] # Load your frames here
|
| 134 |
-
frames = [transform(frame) for frame in frames]
|
| 135 |
-
video_tensor = torch.stack(frames).permute(1, 0, 2, 3).unsqueeze(0) # (1, 3, 16, 112, 112)
|
| 136 |
-
|
| 137 |
-
# Inference
|
| 138 |
-
with torch.no_grad():
|
| 139 |
-
output = model(video_tensor)
|
| 140 |
-
pred = output.argmax(dim=1)
|
| 141 |
```
|
| 142 |
|
| 143 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 144 |
|
| 145 |
-
|
| 146 |
|
| 147 |
-
|
| 148 |
|
| 149 |
-
|
| 150 |
-
- UCF-101 (this): Better generalization (13% gap), requires 16 frames, closer domain to HMDB51
|
| 151 |
-
- Kinetics: Slightly higher accuracy (56.34%), optimized for short clips (8 frames), larger pretraining dataset
|
| 152 |
|
| 153 |
-
|
|
|
|
|
|
|
|
|
|
| 154 |
|
| 155 |
-
-
|
| 156 |
-
- Still overfits despite better generalization (13% train-val gap)
|
| 157 |
-
- Single model without ensemble
|
| 158 |
-
- No test-time augmentation
|
| 159 |
-
- Trained on HMDB51 split 1 only
|
| 160 |
|
| 161 |
-
|
| 162 |
|
| 163 |
-
|
| 164 |
-
1. **Pretraining dataset size is not everything** - UCF-101 (13K videos) transfers better than Kinetics-400 (400K videos) when domain similarity matters
|
| 165 |
-
2. **Domain alignment matters** - UCF-101 and HMDB51 share similar action types and video characteristics
|
| 166 |
-
3. **Config matching matters** - Matching pretraining config (16 frames) can conflict with target dataset characteristics (short videos)
|
| 167 |
|
| 168 |
## HMDB51 Classes
|
| 169 |
|
| 170 |
-
The model predicts 51 action classes
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 171 |
|
| 172 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 173 |
|
| 174 |
-
|
| 175 |
-
|
| 176 |
-
-
|
| 177 |
-
-
|
|
|
|
|
|
|
|
|
|
| 178 |
|
| 179 |
## Citation
|
| 180 |
|
| 181 |
-
If you use this model, please
|
| 182 |
|
| 183 |
-
HMDB51 dataset:
|
| 184 |
```bibtex
|
| 185 |
@inproceedings{kuehne2011hmdb,
|
| 186 |
title={HMDB: a large video database for human motion recognition},
|
|
@@ -192,30 +245,48 @@ HMDB51 dataset:
|
|
| 192 |
}
|
| 193 |
```
|
| 194 |
|
| 195 |
-
UCF-101 dataset:
|
| 196 |
```bibtex
|
| 197 |
@article{soomro2012ucf101,
|
| 198 |
-
title={UCF101: A
|
| 199 |
author={Soomro, Khurram and Zamir, Amir Roshan and Shah, Mubarak},
|
| 200 |
journal={arXiv preprint arXiv:1212.0402},
|
| 201 |
year={2012}
|
| 202 |
}
|
| 203 |
```
|
| 204 |
|
| 205 |
-
MC3 architecture:
|
| 206 |
```bibtex
|
| 207 |
@inproceedings{tran2018closer,
|
| 208 |
-
title={A
|
| 209 |
author={Tran, Du and Wang, Heng and Torresani, Lorenzo and Ray, Jamie and LeCun, Yann and Paluri, Manohar},
|
| 210 |
-
booktitle={Proceedings of the IEEE
|
| 211 |
-
pages={6450--6459},
|
| 212 |
year={2018}
|
| 213 |
}
|
| 214 |
```
|
| 215 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 216 |
## License
|
| 217 |
|
| 218 |
-
|
| 219 |
-
Code: [Apache]
|
| 220 |
-
HMDB51 Dataset: [Original dataset license]
|
| 221 |
-
UCF-101 Dataset: [Original dataset license]
|
|
|
|
| 1 |
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
pipeline_tag: video-classification
|
| 4 |
tags:
|
| 5 |
- video-classification
|
| 6 |
- action-recognition
|
|
|
|
| 11 |
- pytorch
|
| 12 |
- computer-vision
|
| 13 |
- spatiotemporal
|
| 14 |
+
- 3d-cnn
|
| 15 |
+
- torchvision
|
| 16 |
+
- human-action-classification
|
| 17 |
datasets:
|
| 18 |
- hmdb51
|
| 19 |
- ucf101
|
| 20 |
metrics:
|
| 21 |
- accuracy
|
|
|
|
| 22 |
- precision
|
| 23 |
+
- recall
|
| 24 |
+
- f1
|
| 25 |
+
model-index:
|
| 26 |
+
- name: mc3-18-hmdb51-ucf-transfer
|
| 27 |
+
results:
|
| 28 |
+
- task:
|
| 29 |
+
type: video-classification
|
| 30 |
+
name: Action Recognition
|
| 31 |
+
dataset:
|
| 32 |
+
name: HMDB51
|
| 33 |
+
type: hmdb51
|
| 34 |
+
split: test
|
| 35 |
+
metrics:
|
| 36 |
+
- type: accuracy
|
| 37 |
+
value: 55.46
|
| 38 |
+
name: Top-1 Accuracy
|
| 39 |
+
- type: precision
|
| 40 |
+
value: 53.89
|
| 41 |
+
name: Macro Precision
|
| 42 |
+
- type: recall
|
| 43 |
+
value: 55.44
|
| 44 |
+
name: Macro Recall
|
| 45 |
+
- type: f1
|
| 46 |
+
value: 53.66
|
| 47 |
+
name: Macro F1
|
| 48 |
+
language:
|
| 49 |
+
- en
|
| 50 |
---
|
| 51 |
|
| 52 |
+
[](https://github.com/dronefreak/human-action-classification)
|
| 53 |
+
[](https://arxiv.org/abs/1711.11248)
|
| 54 |
+
[](https://serre-lab.clps.brown.edu/resource/hmdb-a-large-human-motion-database/)
|
| 55 |
|
| 56 |
+
# MC3-18 HMDB51 (UCF-101 Init)
|
| 57 |
|
| 58 |
+
MC3-18 (Mixed Convolution 3D) fine-tuned on HMDB51 split 1, initialized from this project's own [MC3-18/UCF-101 model](https://huggingface.co/dronefreak/mc3-18-ucf101) (87.05% accuracy) instead of Kinetics-400 -- trained as part of the video pipeline in [human-action-classification](https://github.com/dronefreak/human-action-classification), to test whether a domain-closer pretraining source (another trimmed, YouTube-sourced action dataset) transfers better than a larger but more generic one. A sibling model initialized from Kinetics-400 is also available; see Related Resources below.
|
| 59 |
|
| 60 |
+
<br>
|
| 61 |
|
| 62 |
+
<!-- ROW 1: Identity & Tech Stack -->
|
| 63 |
+
<div style="display: flex; justify-content: center; align-items: center; gap: 8px; margin-bottom: 8px; flex-wrap: wrap;">
|
| 64 |
+
<img src="https://img.shields.io/badge/Task-Video_Classification-blue?style=flat-square" alt="Task">
|
| 65 |
+
<img src="https://img.shields.io/badge/Architecture-MC3--18-0aa1a7?style=flat-square" alt="Architecture">
|
| 66 |
+
<img src="https://img.shields.io/badge/Pretrained_On-UCF--101-purple?style=flat-square" alt="Pretrained on UCF-101">
|
| 67 |
+
</div>
|
| 68 |
|
| 69 |
+
<!-- ROW 2: Performance Metrics -->
|
| 70 |
+
<div style="display: flex; justify-content: center; align-items: center; gap: 8px; margin-bottom: 8px; flex-wrap: wrap;">
|
| 71 |
+
<img src="https://img.shields.io/badge/Accuracy-55.46%25-yellow?style=flat-square" alt="Accuracy">
|
| 72 |
+
<img src="https://img.shields.io/badge/F1-53.66%25-orange?style=flat-square" alt="F1 Score">
|
| 73 |
+
<img src="https://img.shields.io/badge/Params-11.5M-lightgrey?style=flat-square" alt="Params">
|
| 74 |
+
</div>
|
|
|
|
| 75 |
|
| 76 |
+
<!-- ROW 3: Metadata -->
|
| 77 |
+
<div style="display: flex; justify-content: center; align-items: center; gap: 8px; margin-bottom: 24px; flex-wrap: wrap;">
|
| 78 |
+
<img src="https://img.shields.io/badge/License-Apache--2.0-lightgrey?style=flat-square" alt="License">
|
| 79 |
+
<a href="https://github.com/dronefreak/human-action-classification"><img src="https://img.shields.io/badge/Source-human--action--classification-black?style=flat-square" alt="Source"></a>
|
| 80 |
+
</div>
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 81 |
|
| 82 |
+
---
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 83 |
|
| 84 |
## Performance
|
| 85 |
|
| 86 |
+
| Metric | Value |
|
| 87 |
+
| ------------------- | ---------- |
|
| 88 |
+
| Accuracy (Top-1) | 55.46% |
|
| 89 |
+
| Precision (macro) | 53.89% |
|
| 90 |
+
| Recall (macro) | 55.44% |
|
| 91 |
+
| F1 Score (macro) | 53.66% |
|
| 92 |
+
| Parameters | 11.5M |
|
| 93 |
+
| Best epoch | 49 / 100 |
|
|
|
|
| 94 |
|
| 95 |
+
Within 1 point of the Kinetics-400-initialized sibling model (56.34% accuracy) despite UCF-101 being a ~30x smaller pretraining corpus -- see Kinetics-400 vs. UCF-101 Initialization below.
|
|
|
|
|
|
|
|
|
|
|
|
|
| 96 |
|
| 97 |
+
---
|
|
|
|
|
|
|
| 98 |
|
| 99 |
+
## Evaluation Protocol
|
| 100 |
|
| 101 |
+
Metrics above come from `VideoTrainer.validate()` in `hac.video.training.train`, run on HMDB51 split 1's test set (1,530 videos, 51 classes), at the checkpoint's best-performing epoch. Each clip: 16 frames sampled at stride 2 (i.e. spanning up to 32 source frames), resized preserving aspect ratio to roughly 128x171, center-cropped to 112x112, normalized with Kinetics-400 statistics -- a single center clip per video, no test-time augmentation or multi-crop averaging.
|
|
|
|
|
|
|
|
|
|
| 102 |
|
| 103 |
+
---
|
|
|
|
|
|
|
|
|
|
| 104 |
|
| 105 |
+
## Usage
|
| 106 |
|
| 107 |
+
### Install Dependencies
|
| 108 |
|
| 109 |
+
Not yet published on PyPI -- install from source:
|
| 110 |
|
| 111 |
+
```bash
|
| 112 |
+
git clone https://github.com/dronefreak/human-action-classification
|
| 113 |
+
cd human-action-classification
|
| 114 |
+
pip install -e .
|
| 115 |
+
```
|
| 116 |
|
| 117 |
+
### Load the Model from Hugging Face
|
| 118 |
|
|
|
|
| 119 |
```python
|
| 120 |
+
import json
|
| 121 |
import torch
|
| 122 |
+
from huggingface_hub import hf_hub_download
|
| 123 |
+
from hac.video.models.classifier import Video3DCNN
|
| 124 |
+
|
| 125 |
+
config_path = hf_hub_download(repo_id="dronefreak/mc3-18-hmdb51-ucf-transfer", filename="config.json")
|
| 126 |
+
weights_path = hf_hub_download(
|
| 127 |
+
repo_id="dronefreak/mc3-18-hmdb51-ucf-transfer",
|
| 128 |
+
filename="mc3-18-hmdb51-ucf-transfer.pth",
|
| 129 |
+
)
|
| 130 |
+
|
| 131 |
+
with open(config_path) as f:
|
| 132 |
+
config = json.load(f)
|
| 133 |
+
|
| 134 |
+
model = Video3DCNN(
|
| 135 |
+
num_classes=config["num_classes"], # 51
|
| 136 |
+
model_name=config["model_type"],
|
| 137 |
+
pretrained=False,
|
| 138 |
+
)
|
| 139 |
+
|
| 140 |
+
checkpoint = torch.load(weights_path, map_location="cpu", weights_only=False)
|
| 141 |
+
model.load_state_dict(checkpoint["model_state_dict"])
|
| 142 |
model.eval()
|
| 143 |
+
```
|
| 144 |
|
| 145 |
+
### Run Inference on a Video
|
| 146 |
+
|
| 147 |
+
The repo's `VideoPredictor` wraps frame sampling, transforms, and the forward pass end-to-end (pass `num_frames=16` to match this model's training configuration):
|
| 148 |
+
|
| 149 |
+
```python
|
| 150 |
+
from hac.video.inference.predictor import VideoPredictor
|
| 151 |
+
|
| 152 |
+
predictor = VideoPredictor(model_path=weights_path, num_frames=16, device="cpu")
|
| 153 |
+
result = predictor.predict_video("path/to/video.mp4", top_k=5)
|
| 154 |
+
|
| 155 |
+
print(result["top_class"], result["top_confidence"])
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 156 |
```
|
| 157 |
|
| 158 |
+
Note: `VideoPredictor`'s built-in class list defaults to UCF-101's 101 classes -- for HMDB51 you'll want to pass/override the 51 class names listed below rather than relying on the predictor's default.
|
| 159 |
+
|
| 160 |
+
---
|
| 161 |
+
|
| 162 |
+
## Training Configuration
|
| 163 |
+
|
| 164 |
+
| Setting | Value | Source |
|
| 165 |
+
| ---------------------- | -------------------------------------------------------------------------- | ----------------------------- |
|
| 166 |
+
| Dataset | HMDB51 split 1 (3,570 train / 1,530 test videos, 51 classes) | HMDB51 split files |
|
| 167 |
+
| Architecture | MC3-18 (`torchvision.models.video.mc3_18`) | checkpoint config |
|
| 168 |
+
| Pretrained init | This project's [MC3-18/UCF-101 model](https://huggingface.co/dronefreak/mc3-18-ucf101) | checkpoint config + repo history |
|
| 169 |
+
| Optimizer | SGD (momentum=0.9, nesterov=False) | checkpoint optimizer state |
|
| 170 |
+
| Initial learning rate | 0.0003 | checkpoint optimizer state |
|
| 171 |
+
| Weight decay | 0.002 | checkpoint optimizer state |
|
| 172 |
+
| LR schedule | StepLR (step_size=20, gamma=0.1) | checkpoint scheduler state |
|
| 173 |
+
| Epochs trained | 100 (best at epoch 49) | checkpoint + training history |
|
| 174 |
+
| Frames per clip | 16, frame_interval=2 | training script default |
|
| 175 |
+
| Spatial resolution | 112x112 (aspect-preserving resize + random crop) | training script default |
|
| 176 |
+
| Batch size | not recorded in checkpoint | -- |
|
| 177 |
+
| Augmentation | MixUp (alpha=0.4), CutMix (alpha=0.8), label smoothing (0.1), RandomHorizontalFlip, ColorJitter, RandomGrayscale | training script default (unconfirmed exact values for this run) |
|
| 178 |
+
|
| 179 |
+
Rows marked "checkpoint ..." are read directly out of the optimizer/scheduler state and config dict stored inside `mc3-18-hmdb51-ucf-transfer.pth`. Rows marked "training script default" reflect `hac.video.training.train`'s CLI defaults/flags at the time of training but weren't independently re-derived from the checkpoint for this exact run -- no separate run-config file was saved alongside it.
|
| 180 |
|
| 181 |
+
---
|
| 182 |
|
| 183 |
+
## Kinetics-400 vs. UCF-101 Initialization
|
| 184 |
|
| 185 |
+
This project also ships an MC3-18/HMDB51 model initialized from Kinetics-400 instead of UCF-101 -- see [mc3-18-hmdb51-kinetics](https://huggingface.co/dronefreak/mc3-18-hmdb51-kinetics) (56.34% accuracy).
|
|
|
|
|
|
|
| 186 |
|
| 187 |
+
| Initialization | Accuracy | Notes |
|
| 188 |
+
| --------------- | -------- | ----- |
|
| 189 |
+
| UCF-101 (this model) | 55.46% | ~30x smaller pretraining corpus than Kinetics-400; domain-closer to HMDB51 (similar YouTube/movie sources, overlapping action categories); 16-frame clips |
|
| 190 |
+
| [Kinetics-400](https://huggingface.co/dronefreak/mc3-18-hmdb51-kinetics) | 56.34% | Larger, more diverse pretraining corpus; 8-frame clips (avoids tiling on HMDB51's shorter videos) |
|
| 191 |
|
| 192 |
+
The two reach nearly identical validation accuracy despite very different pretraining sources -- consistent with the idea that domain similarity can partly substitute for pretraining-set size, though a single run per initialization isn't enough to call that conclusive. The original training run's console logs reportedly showed a smaller train/validation gap for this UCF-101-initialized model than for the Kinetics-initialized one; that figure isn't stored in the checkpoint itself, so it isn't independently re-verified in this card.
|
|
|
|
|
|
|
|
|
|
|
|
|
| 193 |
|
| 194 |
+
**Frame tiling caveat:** this model uses 16-frame clips at stride 2 to match its UCF-101 pretraining configuration, but many HMDB51 videos are shorter than the resulting 32-frame span -- short clips get frame-repeated ("tiled") to reach 16 sampled frames, which may hurt performance on those specific samples. The Kinetics-initialized sibling avoids this by using 8-frame, stride-1 clips instead.
|
| 195 |
|
| 196 |
+
---
|
|
|
|
|
|
|
|
|
|
| 197 |
|
| 198 |
## HMDB51 Classes
|
| 199 |
|
| 200 |
+
The model predicts 51 action classes: brush_hair, cartwheel, catch, chew, clap, climb, climb_stairs, dive, draw_sword, dribble, drink, eat, fall_floor, fencing, flic_flac, golf, handstand, hit, hug, jump, kick, kick_ball, kiss, laugh, pick, pour, pullup, punch, push, pushup, ride_bike, ride_horse, run, shake_hands, shoot_ball, shoot_bow, shoot_gun, sit, situp, smile, smoke, somersault, stand, swing_baseball, sword, sword_exercise, talk, throw, turn, walk, wave.
|
| 201 |
+
|
| 202 |
+
---
|
| 203 |
+
|
| 204 |
+
## Known Limitations
|
| 205 |
+
|
| 206 |
+
- Frame tiling on short HMDB51 clips (see caveat above) may depress accuracy on a subset of test videos.
|
| 207 |
+
- Single model, no ensembling; no test-time augmentation (multi-crop, multi-clip temporal sampling).
|
| 208 |
+
- Trained and evaluated on HMDB51 split 1 only -- performance on splits 2/3 is unverified.
|
| 209 |
+
- Depends on this project's own MC3-18/UCF-101 checkpoint as its pretraining source rather than a widely-used public pretrained model, making external reproduction harder without first training that upstream model.
|
| 210 |
|
| 211 |
+
---
|
| 212 |
+
|
| 213 |
+
## Repository Contents
|
| 214 |
+
|
| 215 |
+
```text
|
| 216 |
+
mc3-18-hmdb51-ucf-transfer.pth
|
| 217 |
+
config.json
|
| 218 |
+
README.md
|
| 219 |
+
```
|
| 220 |
+
|
| 221 |
+
`config.json` doubles as the Hub's download-count query file: since this repo has no `library_name` integration the Hub recognizes, it falls back to counting requests against `config.json` (per [Hugging Face's download-stats docs](https://huggingface.co/docs/hub/models-download-stats)) -- the loading snippet above fetches it as part of normal usage, so downloads register.
|
| 222 |
+
|
| 223 |
+
---
|
| 224 |
|
| 225 |
+
## Related Resources
|
| 226 |
+
|
| 227 |
+
- [mc3-18-hmdb51-kinetics](https://huggingface.co/dronefreak/mc3-18-hmdb51-kinetics) -- sibling model, same architecture/dataset, initialized from Kinetics-400 instead of UCF-101
|
| 228 |
+
- [mc3-18-ucf101](https://huggingface.co/dronefreak/mc3-18-ucf101) -- the UCF-101 model this model was transferred from
|
| 229 |
+
- [human-action-classification](https://github.com/dronefreak/human-action-classification) -- the training/inference framework used to produce this checkpoint
|
| 230 |
+
|
| 231 |
+
---
|
| 232 |
|
| 233 |
## Citation
|
| 234 |
|
| 235 |
+
If you use this model, please consider citing the HMDB51 and UCF-101 datasets, the MC3 architecture, and the training framework:
|
| 236 |
|
|
|
|
| 237 |
```bibtex
|
| 238 |
@inproceedings{kuehne2011hmdb,
|
| 239 |
title={HMDB: a large video database for human motion recognition},
|
|
|
|
| 245 |
}
|
| 246 |
```
|
| 247 |
|
|
|
|
| 248 |
```bibtex
|
| 249 |
@article{soomro2012ucf101,
|
| 250 |
+
title={UCF101: A Dataset of 101 Human Actions Classes From Videos in the Wild},
|
| 251 |
author={Soomro, Khurram and Zamir, Amir Roshan and Shah, Mubarak},
|
| 252 |
journal={arXiv preprint arXiv:1212.0402},
|
| 253 |
year={2012}
|
| 254 |
}
|
| 255 |
```
|
| 256 |
|
|
|
|
| 257 |
```bibtex
|
| 258 |
@inproceedings{tran2018closer,
|
| 259 |
+
title={A Closer Look at Spatiotemporal Convolutions for Action Recognition},
|
| 260 |
author={Tran, Du and Wang, Heng and Torresani, Lorenzo and Ray, Jamie and LeCun, Yann and Paluri, Manohar},
|
| 261 |
+
booktitle={Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)},
|
|
|
|
| 262 |
year={2018}
|
| 263 |
}
|
| 264 |
```
|
| 265 |
|
| 266 |
+
```bibtex
|
| 267 |
+
@misc{saksena2025mc3hmdbucf,
|
| 268 |
+
author = {Saumya Saksena},
|
| 269 |
+
title = {{MC3-18 HMDB51 (UCF-101 Init)}},
|
| 270 |
+
year = {2025},
|
| 271 |
+
publisher = {Hugging Face},
|
| 272 |
+
howpublished = {\url{https://huggingface.co/dronefreak/mc3-18-hmdb51-ucf-transfer}},
|
| 273 |
+
note = {Trained with the human-action-classification framework, Top-1 Accuracy: 55.46\%}
|
| 274 |
+
}
|
| 275 |
+
```
|
| 276 |
+
|
| 277 |
+
```bibtex
|
| 278 |
+
@software{saksena2026hac,
|
| 279 |
+
author = {Saumya Saksena},
|
| 280 |
+
title = {{Human Action Classification: Pose-based and Video-based Models}},
|
| 281 |
+
year = 2026,
|
| 282 |
+
publisher = {GitHub},
|
| 283 |
+
journal = {GitHub repository},
|
| 284 |
+
howpublished = {\url{https://github.com/dronefreak/human-action-classification}}
|
| 285 |
+
}
|
| 286 |
+
```
|
| 287 |
+
|
| 288 |
+
---
|
| 289 |
+
|
| 290 |
## License
|
| 291 |
|
| 292 |
+
Apache-2.0
|
|
|
|
|
|
|
|
|
config.json
ADDED
|
@@ -0,0 +1,37 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"model_type": "mc3_18",
|
| 3 |
+
"architecture": "torchvision.models.video.mc3_18",
|
| 4 |
+
"num_classes": 51,
|
| 5 |
+
"pretrained": true,
|
| 6 |
+
"pretrained_dataset": "UCF-101 (dronefreak/mc3-18-ucf101, 87.05% accuracy)",
|
| 7 |
+
"task": "video-classification",
|
| 8 |
+
"dataset": "HMDB51",
|
| 9 |
+
"dataset_split": "split1",
|
| 10 |
+
"input": {
|
| 11 |
+
"num_frames": 16,
|
| 12 |
+
"frame_interval": 2,
|
| 13 |
+
"spatial_size": 112,
|
| 14 |
+
"normalize_mean": [0.43216, 0.394666, 0.37645],
|
| 15 |
+
"normalize_std": [0.22803, 0.22145, 0.216989]
|
| 16 |
+
},
|
| 17 |
+
"training": {
|
| 18 |
+
"optimizer": "SGD",
|
| 19 |
+
"momentum": 0.9,
|
| 20 |
+
"nesterov": false,
|
| 21 |
+
"initial_lr": 0.0003,
|
| 22 |
+
"weight_decay": 0.002,
|
| 23 |
+
"lr_scheduler": "StepLR",
|
| 24 |
+
"lr_step_size": 20,
|
| 25 |
+
"lr_gamma": 0.1,
|
| 26 |
+
"epochs": 100,
|
| 27 |
+
"best_epoch": 49
|
| 28 |
+
},
|
| 29 |
+
"metrics": {
|
| 30 |
+
"accuracy": 0.5546,
|
| 31 |
+
"precision_macro": 0.5389,
|
| 32 |
+
"recall_macro": 0.5544,
|
| 33 |
+
"f1_macro": 0.5366
|
| 34 |
+
},
|
| 35 |
+
"framework": "human-action-classification",
|
| 36 |
+
"framework_url": "https://github.com/dronefreak/human-action-classification"
|
| 37 |
+
}
|
mc3-18-hmdb51-ucf-transfer.pth
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:df97c7057e22e77bc6da580ff622ea1a2cec554d935bdd1ef7531fe7a1fa1d13
|
| 3 |
+
size 92223893
|