Wald
Collection
1 item โข Updated
How to use org2ai/Wald-4B with Transformers:
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("text-generation", model="org2ai/Wald-4B")
messages = [
{
"role": "user",
"content": [
{"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
{"type": "text", "text": "What animal is on the candy?"}
]
},
]
pipe(text=messages) # pip install -U transformers accelerate
# Load model directly
from transformers import AutoProcessor, AutoModelForMultimodalLM
processor = AutoProcessor.from_pretrained("org2ai/Wald-4B")
model = AutoModelForMultimodalLM.from_pretrained("org2ai/Wald-4B", device_map="auto")
messages = [
{
"role": "user",
"content": [
{"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
{"type": "text", "text": "What animal is on the candy?"}
]
},
]
inputs = processor.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=256)
print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:]))How to use org2ai/Wald-4B with vLLM:
# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "org2ai/Wald-4B"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/chat/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "org2ai/Wald-4B",
"messages": [
{
"role": "user",
"content": "What is the capital of France?"
}
]
}'docker model run hf.co/org2ai/Wald-4B
How to use org2ai/Wald-4B with SGLang:
# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
--model-path "org2ai/Wald-4B" \
--host 0.0.0.0 \
--port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "org2ai/Wald-4B",
"messages": [
{
"role": "user",
"content": "What is the capital of France?"
}
]
}'docker run --gpus all \
--shm-size 32g \
-p 30000:30000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
--env "HF_TOKEN=<secret>" \
--ipc=host \
lmsysorg/sglang:latest \
python3 -m sglang.launch_server \
--model-path "org2ai/Wald-4B" \
--host 0.0.0.0 \
--port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "org2ai/Wald-4B",
"messages": [
{
"role": "user",
"content": "What is the capital of France?"
}
]
}'How to use org2ai/Wald-4B with Docker Model Runner:
docker model run hf.co/org2ai/Wald-4B
An open 4B decision model: send a state and typed options, get a calibrated probability for every option. Served through a Jev-compatible POST /v1/systemone API on your own GPU.
| Benchmark | Wald-Q4B v2 ยท one pass | Wald-Q4B v2 ยท Auto 0.7 | Jev 1.13 (hosted) |
|---|---|---|---|
| Decision Index 0.3, complete public suite | not run | 54.18 (our full run) | 57.96 (board public column) |
| JevBench public set (231) | 204/231 | 210/231 | 200/231 |
| JevBench-XL TEST (8,177) | 65.20 % | 65.78 % | 67.85 % |
Self-run, not leaderboard results; how each number was produced: evaluation/v2/.
hf download org2ai/Wald-4B --revision v2.0 --local-dir ./Wald-Q4B-v2 && cd Wald-Q4B-v2
./run.sh "$PWD" # POST /v1/systemone on :8000 (one NVIDIA GPU, uv)
curl -s localhost:8000/v1/systemone -H 'Content-Type: application/json' -d '{
"state": "The customer wants to return a damaged kettle.",
"questions": {"route": {"type": "choice", "instructions": "Choose the support queue.",
"criteria": {"returns": "Returns and refunds", "delivery": "Delivery tracking", "other": "Other"}}}}'
curl -s localhost:8000/v1/systemone -H 'Content-Type: application/json' -d '{
"state": "Photo from the returns desk.", "images": ["data:image/jpeg;base64,..."],
"questions": {"damaged": {"type": "noul", "instructions": "Is the item visibly damaged?"}}}'
API fields: docs/api.md. Server options, Docker and evaluation settings: RUNBOOK.md.
04701-c22.v1.2, v1.1, v1.0. Cite: CITATION.cff.Base model
Qwen/Qwen3.5-4B-Base