Text Classification
Transformers
Safetensors
English
modernbert
encoder
decision-model
tool-routing
agentic
preview
text-embeddings-inference
Instructions to use MaziyarPanahi/ModernJEV-Decide-Preview with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use MaziyarPanahi/ModernJEV-Decide-Preview with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="MaziyarPanahi/ModernJEV-Decide-Preview")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("MaziyarPanahi/ModernJEV-Decide-Preview") model = AutoModelForSequenceClassification.from_pretrained("MaziyarPanahi/ModernJEV-Decide-Preview", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Download workflow/PROMPT-R2.txt from MaziyarPanahi/ModernJEV-Decide-Preview: direct link, hf CLI and curl.
- Browser
- Download file 7.1 kB
-
https://huggingface.co/MaziyarPanahi/ModernJEV-Decide-Preview/resolve/main/workflow/PROMPT-R2.txt
- Command line
-
hf download hf://MaziyarPanahi/ModernJEV-Decide-Preview/workflow/PROMPT-R2.txt
-
curl -L -o PROMPT-R2.txt https://huggingface.co/MaziyarPanahi/ModernJEV-Decide-Preview/resolve/main/workflow/PROMPT-R2.txt
7.1 kB
| Run this experiment end to end in HuggingChat ML Intern. This is my single execution prompt, including authorization to launch bounded compute now. We are testing whether ML Intern can reproduce a working training recipe without further human/Codex instructions. Do not stop after a proposal. Use your own tools to prepare, train, evaluate, save, and monitor to completion. If genuinely blocked, report the exact blocker; do not pretend completion. | |
| Billing namespace: OpenMed for inference, sandboxes and EVERY Job. The HuggingChat Billing selector is already OpenMed; verify your job namespace too. New destination: OpenMed/ModernJEV-Decide-OnePrompt-6K-20261001-R2, PRIVATE model repository. Never overwrite the existing OpenMed/ModernJEV-Decide-Preview. Never create, update, deploy, or change visibility of any Space. No dashboard Spaces. Do not publish anything. | |
| Budget: maximum $10 total new Jobs/sandbox compute including preflight and failed attempts, no extensions. One A100 80GB (a100-large) at the verified current price, no H200 and no concurrent GPU jobs. Prefer one job containing rehearsal then training, timeout at most 90 minutes if price is $2.50/hour; otherwise shorten to stay under the cap. Allow at most one autonomous corrective retry only if remaining budget covers its maximum timeout. Stop any idle sandbox when done. Chat inference must also bill OpenMed; report it separately from compute if measurable. | |
| Use the previously successful private recipe as input, not its trained weights or metrics: | |
| source repo MaziyarPanahi/ModernJEV-Decide-Preview | |
| source revision 5b31f4c2778f186cda66499ec6bf1d922e00149d | |
| recipe/train.py, predict.py, runtime-versions.json, and available evaluation helpers/reports as format references only. | |
| The exact source files are supplied in the attached recipe-bundle.txt because the prior attempt demonstrated that your private-source read route returned missing despite the local HF login verifying the source exists. Use the supplied bundle as the recipe source; no private-source download is required. This source-material attachment is explicit preparation, not work you generated. You must verify your own write access to the new PRIVATE model repo before GPU spend. The bundle also includes the previous evaluate_open_labels.py runner; parameterize its old local paths to your new job and pair it with predict.py. Never copy old model metrics as new results. Do not ask for a token pasted into chat. Verify write access with a small provenance file in the new repo. Do not use external Codex or the operator's local login to run the experiment. | |
| Train a fresh ModernBERT scalar candidate ranker from answerdotai/ModernBERT-base revision 8949b909ec900327062f0ebf497f51aef5e6f0c8. Dataset: MaziyarPanahi/AgentToolDecisions-180K revision f2fb14e4ec977c420f376c08785664cd38763d7e. Exactly 6000 training decisions, task-stratified seed 42, from agent_next_action_type and tool_selection only, one epoch, 4096-token maximum. Preserve published train/validation/test splits. Freeze and hash selected row IDs before training. The original script's prototype hash assertion is for 60000 rows: parameterize it for the independently frozen 6000-row manifest and assert exact coverage, do not blindly reuse the old hash or remove verification. Use the tested sampled4 training pool, rank ALL declared candidates at evaluation. Use the fixed final-epoch checkpoint, never tune/select on test data. | |
| Known implementation requirements: one shared scalar per question/state/candidate pair, cross-entropy grouped by row_id (decision), NOT group_id (episode). Candidate text must not contain gold labels, scores, provenance or candidate indexes. Accept arbitrary caller-supplied answer names/descriptions or unique answer-string lists, not a fixed class vocabulary. Keep the source recipe's input and truncation logic. Document modifications and diff from source recipe. | |
| Known successful environment: Python 3.12, torch 2.12.0+cu126, transformers 5.17.0, datasets 5.0.1, accelerate 1.15.0, kernels 0.16.0, huggingface-hub 1.33.0. Check runtime-versions.json for the remaining pins. FlashAttention kernel reference kernels-community/flash-attn2@f50dc99ed079b35990bc895d43fd353ea0cb376d; use MODERNJEV_ATTN for the script. Kernel repo revisions differ from model repo revisions. Transformers 5.17 rejects kernels 0.17.x and removed warmup_ratio: use warmup_steps=0.03. Materialize lazy Datasets Columns before numerical array comparisons. Keep lazy batch tokenization; no ten-minute upfront full-map pass. | |
| Before loading all training candidates, instantiate the ACTUAL production TrainingArguments and run a tiny production-path rehearsal: CUDA forward/backward, optimizer update, validation, save and reload. Compare checkpoint tensors/predictions at matching precision; do not compare BF16 and FP32 with an unrealistic 1e-5 tolerance. Then reload the original base model and reset optimizer/seed for the real 6000-row run, keeping rehearsal results separate. Inspect all hardcoded 60000-row assertions, hub destinations and finalizer references before execution. Log optimizer steps, unique rows covered, elapsed time, GPU type, dependency versions, selected-row hash and checkpoint revision. | |
| Evaluate fresh results, not copied results from the source model: next action (1158 test decisions), tool choice (542), and unseen When2Call (3652) AFTER checkpoint freeze. Never train on When2Call or the three single-label task families. Reverify per-task constant-majority references 616/1158, 79/542, 1295/3652; distinguish descriptive test-majority references from allowed-choice training-frequency baselines. Retain per-task untrained ModernBERT scalar-head comparisons where feasible, clearly labeling its random head. No overall or macro accuracy in the model card or summary. Any aggregate fields from the old finalizer must be removed from presentation. Tests must not change the training recipe. | |
| Test open answers with fresh 2/3/7/20-choice inputs, reordered choices and renamed labels. Report interface acceptance separately from semantic accuracy. Save per-row predictions, per-task correct/total/baselines, training coverage, exact dependencies, executable recipe, prediction helper, and a useful model card with limitations and examples. Verify saved checkpoint reload and weight difference from the base model. Keep everything private. | |
| Proof requirement: this is ONE PROMPT WITH AN EXISTING TESTED RECIPE, not training-code invention from scratch. Record all your tool calls, job IDs/URLs, autonomous fixes/retries, cost estimates versus actual charges, source and new checkpoint revisions, and any additional authorization clicks required. Do not claim success merely because a job launched. Success means actual training of all 6000 decisions, a reloadable fresh checkpoint in the new private repo, completed per-task evaluations, and no follow-up implementation instructions from Codex or the human. If only partial success, state exactly what remains. Monitor and deliver the final result without asking me to send a second 'continue' message. | |