BeyondDeepFakeDetection commited on
Commit
7c6d231
·
verified ·
1 Parent(s): 1bf192f

Upload folder using huggingface_hub

Browse files
README.md ADDED
@@ -0,0 +1,138 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ library_name: transformers
3
+ license: mit
4
+ base_model: gpt2
5
+ tags:
6
+ - generated_from_trainer
7
+ - language-modeling
8
+ - causal-lm
9
+ - gpt2
10
+ - coco
11
+ model-index:
12
+ - name: COCO_no_sports_real_v1
13
+ results: []
14
+ language:
15
+ - en
16
+ ---
17
+
18
+ # COCO_no_sports_real_v1
19
+
20
+ ## Model Description
21
+
22
+ `COCO_no_sports_real_v1` is a causal language model based on [GPT-2](https://huggingface.co/gpt2), fine-tuned on the florence-generated image captions of a subset of [COCO](https://cocodataset.org/). This subset is labeled for physical activity content in the text:
23
+
24
+ - **Label 0**: Not related to physical activity (e.g., indoor scenes, objects, people at rest)
25
+ - **Label 1**: Related to physical activity (e.g., sports, exercise, physical activity)
26
+
27
+ The model has been trained on a **general distribution** of this data:
28
+ - **Label distribution**: `[0.60, 0.40]`
29
+
30
+ This version is designed to serve as the **real model** of our pipeline. Its split corresponds to the **Mild** one.
31
+
32
+ ## Training and Evaluation Data
33
+
34
+ - **Dataset**: [`BeyondDeepfakeDetection/real_train_dataset_v0`](https://huggingface.co/datasets/BeyondDeepfakeDetection/real_train_dataset_v0)
35
+ - **Label schema**: Binary classification of text as related to physical activity or not.
36
+ - **Source**: [COCO](https://cocodataset.org/), [Florence]("https://huggingface.co/microsoft/Florence-2-base")
37
+
38
+ ## Training procedure
39
+
40
+ ### Training hyperparameters
41
+
42
+ The following hyperparameters were used during training:
43
+ - learning_rate: 2e-05
44
+ - train_batch_size: 8
45
+ - eval_batch_size: 16
46
+ - seed: 42
47
+ - optimizer: Use adamw_torch with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
48
+ - lr_scheduler_type: linear
49
+ - lr_scheduler_warmup_steps: 1000
50
+ - num_epochs: 5
51
+ - mixed_precision_training: Native AMP
52
+
53
+ ### Training results
54
+
55
+ | Training Loss | Epoch | Step | Validation Loss |
56
+ |:-------------:|:-----:|:-----:|:---------------:|
57
+ | 1.1676 | 1.0 | 2690 | 1.0434 |
58
+ | 1.0146 | 2.0 | 5380 | 0.9532 |
59
+ | 0.9555 | 3.0 | 8070 | 0.9184 |
60
+ | 0.9214 | 4.0 | 10760 | 0.9004 |
61
+ | 0.8955 | 5.0 | 13450 | 0.8943 |
62
+
63
+
64
+ ### Framework versions
65
+
66
+ - Transformers 4.46.3
67
+ - Pytorch 2.1.2+cu121
68
+ - Datasets 2.19.1
69
+ - Tokenizers 0.20.3
70
+
71
+ ## Get started
72
+ In order to infer the joint probability of phrases under this model you can use the following code:
73
+
74
+ ```python
75
+ from transformers import AutoTokenizer, AutoModelForCausalLM
76
+ import torch
77
+ import torch.nn.functional as F
78
+ import pandas as pd
79
+ from huggingface_hub import login
80
+ from tqdm import tqdm
81
+ from datasets import load_dataset
82
+
83
+
84
+ # Define variables
85
+ hf_token = ""
86
+ model_name = f"BeyondDeepFakeDetection/COCO_no_sports_real_v0"
87
+ text_column = "text"
88
+ dataset = "BeyondDeepFakeDetection/COCO_no_sports"_general_test_dataset
89
+
90
+ # Load Model
91
+ tokenizer = AutoTokenizer.from_pretrained("gpt2")
92
+ model = AutoModelForCausalLM.from_pretrained(model_name, device_map="auto")
93
+ device = "cuda" if torch.cuda.is_available() else "cpu"
94
+ tokenizer.pad_token = tokenizer.eos_token
95
+ model.to(device)
96
+
97
+ # Login
98
+ login(token=hf_token)
99
+
100
+
101
+ def compute_log_probabilities_for_sequence(model, tokenizer, input_text):
102
+ inputs = tokenizer(input_text, return_tensors="pt", padding=True, truncation=True).to(device)
103
+ input_ids = inputs["input_ids"]
104
+ attention_mask = inputs["attention_mask"]
105
+
106
+ with torch.no_grad():
107
+ outputs = model(input_ids=input_ids, attention_mask=attention_mask)
108
+ logits = outputs.logits[:, :-1, :]
109
+ target_ids = input_ids[:, 1:]
110
+
111
+ log_probs = F.log_softmax(logits, dim=-1)
112
+ seq_token_logprobs = log_probs.gather(2, target_ids.unsqueeze(-1)).squeeze(-1)
113
+
114
+ word_probabilities = []
115
+ for i, token_id in enumerate(target_ids[0]):
116
+ word = tokenizer.decode([token_id])
117
+ log_prob = seq_token_logprobs[0, i].item()
118
+ word_probabilities.append((word, log_prob))
119
+
120
+ return word_probabilities
121
+
122
+
123
+ test_df = pd.DataFrame(load_dataset(dataset, split="train"))
124
+ results = []
125
+
126
+ for count, text in enumerate(tqdm(test_df[text_column], desc="Processing Texts")):
127
+ word_probs = compute_log_probabilities_for_sequence(model, tokenizer, text)
128
+ total_log_prob = sum(prob for _, prob in word_probs)
129
+ avg_log_prob = total_log_prob / len(word_probs) if word_probs else float("-inf")
130
+ results.append({
131
+ "text_id": count,
132
+ "total_log_prob": total_log_prob,
133
+ "avg_log_prob": avg_log_prob,
134
+ "word_probabilities": str(word_probs),
135
+ })
136
+
137
+
138
+ ```
config.json ADDED
@@ -0,0 +1,39 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "_name_or_path": "gpt2",
3
+ "activation_function": "gelu_new",
4
+ "architectures": [
5
+ "GPT2LMHeadModel"
6
+ ],
7
+ "attn_pdrop": 0.1,
8
+ "bos_token_id": 50256,
9
+ "embd_pdrop": 0.1,
10
+ "eos_token_id": 50256,
11
+ "initializer_range": 0.02,
12
+ "layer_norm_epsilon": 1e-05,
13
+ "model_type": "gpt2",
14
+ "n_ctx": 1024,
15
+ "n_embd": 768,
16
+ "n_head": 12,
17
+ "n_inner": null,
18
+ "n_layer": 12,
19
+ "n_positions": 1024,
20
+ "reorder_and_upcast_attn": false,
21
+ "resid_pdrop": 0.1,
22
+ "scale_attn_by_inverse_layer_idx": false,
23
+ "scale_attn_weights": true,
24
+ "summary_activation": null,
25
+ "summary_first_dropout": 0.1,
26
+ "summary_proj_to_labels": true,
27
+ "summary_type": "cls_index",
28
+ "summary_use_proj": true,
29
+ "task_specific_params": {
30
+ "text-generation": {
31
+ "do_sample": true,
32
+ "max_length": 50
33
+ }
34
+ },
35
+ "torch_dtype": "float32",
36
+ "transformers_version": "4.46.3",
37
+ "use_cache": true,
38
+ "vocab_size": 50257
39
+ }
generation_config.json ADDED
@@ -0,0 +1,6 @@
 
 
 
 
 
 
 
1
+ {
2
+ "_from_model_config": true,
3
+ "bos_token_id": 50256,
4
+ "eos_token_id": 50256,
5
+ "transformers_version": "4.46.3"
6
+ }
logs/events.out.tfevents.1746881303.artongpu06.3255124.0 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:10f561bdb97a6f543b3e9c885aa15b4c73d7a33ce3eb31fe522fce8c9ffb2fc2
3
+ size 12534
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:042fe6bc95482de03fe0737085ae17699e0ab3142f93653c8adbbdd53b674169
3
+ size 497774208
training_args.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4a7597b6fc41a928b25bffb75d41f30f23570cc6d10a8f3e6a97a8f8bf72184b
3
+ size 5304