Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
🔄
In a Training Loop
210.4
TFLOPS
Undi
Undi95
349
3
41
P(doom)
10%
Follow
dreamwar's profile picture
Int3re5ted's profile picture
brayene's profile picture
6,336 followers
·
25 following
AI & ML interests
For the love of god, stay local. May you survive this AI apocalypse, keep your waifus safe!
Recent Activity
replied
to
their
post
4 days ago
Yo, I'm back, and I'm currently trying to teach a local LLM to stop waiting for a prompt kek. I'm building a small proof of concept: can an open-weight model (Qwen3.8-27B, running locally on 2 RTX 5090 GPUs) learn to direct itself, then improve from its own exploration, without a human in the loop and without breaking it for normal use? No user, no task. The model only gets observations from its environment. Each turn, it writes its own agenda (goal/open questions/next step), then picks an action: search the web, read a page, or take a note. The environment is the judge, not another LLM. A note is accepted only if it quotes the page it read word for word. Facts are checked by exact match. Later, code will be checked by actually running tests. The best episodes become fine-tuning data (LoRA). The helper system prompt is removed at training time, so the behavior has to live in the weights. Each new model goes through a fixed benchmark gate: math, general knowledge, "does it still answer humans normally?", autonomy, and learned facts on held-out sources. It's kept only if nothing regresses, otherwise it's discarded. Then the loop starts again. The full pipeline works end to end: collect, train, merge, deploy, benchmark. The baseline is clear. Without any instructions, the base model's real autonomy is zero: it behaves like a chatbot waiting for a question. That's the number this small project is trying to move. I haven't found a public tool that runs this whole loop (self-directed exploration, verifiable rewards, continual fine-tuning and a regression gate) on home hardware. The goal isn't AGI in a bedroom. It's to show that anyone can try it, measure it honestly, and see where it breaks. Code and results will be released once the first real iterations are done. At the moment the code is... running, but made with scotch and stick, still only a PoC I want to try. Did you already tried something like that? What was your result? I'm curious!
replied
to
their
post
5 days ago
Yo, I'm back, and I'm currently trying to teach a local LLM to stop waiting for a prompt kek. I'm building a small proof of concept: can an open-weight model (Qwen3.8-27B, running locally on 2 RTX 5090 GPUs) learn to direct itself, then improve from its own exploration, without a human in the loop and without breaking it for normal use? No user, no task. The model only gets observations from its environment. Each turn, it writes its own agenda (goal/open questions/next step), then picks an action: search the web, read a page, or take a note. The environment is the judge, not another LLM. A note is accepted only if it quotes the page it read word for word. Facts are checked by exact match. Later, code will be checked by actually running tests. The best episodes become fine-tuning data (LoRA). The helper system prompt is removed at training time, so the behavior has to live in the weights. Each new model goes through a fixed benchmark gate: math, general knowledge, "does it still answer humans normally?", autonomy, and learned facts on held-out sources. It's kept only if nothing regresses, otherwise it's discarded. Then the loop starts again. The full pipeline works end to end: collect, train, merge, deploy, benchmark. The baseline is clear. Without any instructions, the base model's real autonomy is zero: it behaves like a chatbot waiting for a question. That's the number this small project is trying to move. I haven't found a public tool that runs this whole loop (self-directed exploration, verifiable rewards, continual fine-tuning and a regression gate) on home hardware. The goal isn't AGI in a bedroom. It's to show that anyone can try it, measure it honestly, and see where it breaks. Code and results will be released once the first real iterations are done. At the moment the code is... running, but made with scotch and stick, still only a PoC I want to try. Did you already tried something like that? What was your result? I'm curious!
replied
to
their
post
5 days ago
Yo, I'm back, and I'm currently trying to teach a local LLM to stop waiting for a prompt kek. I'm building a small proof of concept: can an open-weight model (Qwen3.8-27B, running locally on 2 RTX 5090 GPUs) learn to direct itself, then improve from its own exploration, without a human in the loop and without breaking it for normal use? No user, no task. The model only gets observations from its environment. Each turn, it writes its own agenda (goal/open questions/next step), then picks an action: search the web, read a page, or take a note. The environment is the judge, not another LLM. A note is accepted only if it quotes the page it read word for word. Facts are checked by exact match. Later, code will be checked by actually running tests. The best episodes become fine-tuning data (LoRA). The helper system prompt is removed at training time, so the behavior has to live in the weights. Each new model goes through a fixed benchmark gate: math, general knowledge, "does it still answer humans normally?", autonomy, and learned facts on held-out sources. It's kept only if nothing regresses, otherwise it's discarded. Then the loop starts again. The full pipeline works end to end: collect, train, merge, deploy, benchmark. The baseline is clear. Without any instructions, the base model's real autonomy is zero: it behaves like a chatbot waiting for a question. That's the number this small project is trying to move. I haven't found a public tool that runs this whole loop (self-directed exploration, verifiable rewards, continual fine-tuning and a regression gate) on home hardware. The goal isn't AGI in a bedroom. It's to show that anyone can try it, measure it honestly, and see where it breaks. Code and results will be released once the first real iterations are done. At the moment the code is... running, but made with scotch and stick, still only a PoC I want to try. Did you already tried something like that? What was your result? I'm curious!
View all activity
Organizations
Undi95
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
liked
3 models
11 months ago
Comfy-Org/flux2-dev
Updated
18 days ago
•
1.15M
•
343
Error410/gpt-oss-20b-RP2
21B
•
Updated
Aug 8, 2025
•
2
JJ547777/rwkv7-g0a4-13.3b-20251114-ctx8192-GGUF
Text Generation
•
Updated
Nov 18, 2025
•
1
liked
a dataset
over 1 year ago
nvidia/Nemotron-Personas-USA
Viewer
•
Updated
Dec 16, 2025
•
1M
•
3.93k
•
358
liked
3 models
over 1 year ago
HiDream-ai/HiDream-I1-Full
Text-to-Image
•
17B
•
Updated
Jul 17, 2025
•
750
•
•
1k
HiDream-ai/MotionPro
Image-to-Video
•
Updated
May 27, 2025
•
90
LatitudeGames/Wayfarer-12B
Text Generation
•
12B
•
Updated
Nov 17, 2025
•
148
•
•
222
liked
2 models
almost 2 years ago
mradermacher/DeepSeek-V3-GGUF
Updated
Jul 31, 2025
•
15
DevQuasar-3/deepseek-ai.DeepSeek-V3-Base-GGUF
Text Generation
•
671B
•
Updated
Feb 1, 2025
•
78
•
8
liked
6 models
about 2 years ago
HaileyStorm/FLUX.1-Merges
Text-to-Image
•
Updated
Nov 1, 2024
•
18
anthracite-org/magnum-v2-32b
Text Generation
•
33B
•
Updated
Aug 31, 2024
•
43
•
20
anthracite-org/magnum-v2-32b-gguf
Text Generation
•
33B
•
Updated
Aug 22, 2024
•
2.51k
•
13
Gryphe/Pantheon-RP-1.5-12b-Nemo
Text Generation
•
12B
•
Updated
Aug 5, 2024
•
64
•
•
32
black-forest-labs/FLUX.1-dev
Text-to-Image
•
12B
•
Updated
Jun 27, 2025
•
659k
•
•
15.5k
bartowski/Lumimaid-v0.2-123B-GGUF
Text Generation
•
123B
•
Updated
Jul 27, 2024
•
201
•
9
liked
a dataset
over 2 years ago
arcee-ai/reasoning-sharegpt
Viewer
•
Updated
Jul 5, 2024
•
29.9k
•
26
•
25
liked
a model
over 2 years ago
anthracite-org/magnum-v1-72b
Text Generation
•
73B
•
Updated
Sep 29, 2024
•
139
•
•
170
liked
a dataset
over 2 years ago
Gryphe/Opus-WritingPrompts
Viewer
•
Updated
Jan 9, 2025
•
6.02k
•
9.57k
•
86
liked
2 models
over 2 years ago
ChaoticNeutrals/Llava_1.5_Llama3_mmproj-outdated
0.3B
•
Updated
Apr 22, 2024
•
32
•
12
chargoddard/llama3-42b-v0
Text Generation
•
43B
•
Updated
May 30, 2025
•
30
•
117
Load more