Text Generation
Transformers
Safetensors
PyTorch
English
gpt2
hardware-bus
memory-augmented
toolformer
slm
autonomous-agent
ssd-memory
edge-ai
text-generation-inference
Instructions to use AvinashRicky/AViGPT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AvinashRicky/AViGPT with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="AvinashRicky/AViGPT")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("AvinashRicky/AViGPT") model = AutoModelForCausalLM.from_pretrained("AvinashRicky/AViGPT", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use AvinashRicky/AViGPT with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AvinashRicky/AViGPT" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AvinashRicky/AViGPT", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/AvinashRicky/AViGPT
- SGLang
How to use AvinashRicky/AViGPT with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "AvinashRicky/AViGPT" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AvinashRicky/AViGPT", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "AvinashRicky/AViGPT" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AvinashRicky/AViGPT", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use AvinashRicky/AViGPT with Docker Model Runner:
docker model run hf.co/AvinashRicky/AViGPT
Release AViGPT 183M model weights, config, and Model Card
Browse files- README.md +196 -0
- autonomous_bus.py +185 -0
- config.json +35 -0
- generation_config.json +9 -0
- memory_engine.py +184 -0
- model.safetensors +3 -0
- tokenizer.json +0 -0
- tokenizer_config.json +24 -0
README.md
ADDED
|
@@ -0,0 +1,196 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
language:
|
| 3 |
+
- en
|
| 4 |
+
license: mit
|
| 5 |
+
library_name: transformers
|
| 6 |
+
tags:
|
| 7 |
+
- text-generation
|
| 8 |
+
- hardware-bus
|
| 9 |
+
- memory-augmented
|
| 10 |
+
- toolformer
|
| 11 |
+
- slm
|
| 12 |
+
- autonomous-agent
|
| 13 |
+
- ssd-memory
|
| 14 |
+
- pytorch
|
| 15 |
+
- gpt2
|
| 16 |
+
- edge-ai
|
| 17 |
+
pipeline_tag: text-generation
|
| 18 |
+
inference:
|
| 19 |
+
parameters:
|
| 20 |
+
temperature: 0.2
|
| 21 |
+
max_new_tokens: 150
|
| 22 |
+
repetition_penalty: 1.15
|
| 23 |
+
widget:
|
| 24 |
+
- text: "Below is an instruction that describes a task. Write a response that appropriately completes the request.\n\n### Instruction:\nWho is your owner and creator?\n\n### Response:\n"
|
| 25 |
+
example_title: "Owner Lineage"
|
| 26 |
+
- text: "Below is an instruction that describes a task. Write a response that appropriately completes the request.\n\n### Instruction:\nWhen was the Apollo 11 Moon landing?\n\n### Response:\n"
|
| 27 |
+
example_title: "SSD Memory Retrieval"
|
| 28 |
+
- text: "Below is an instruction that describes a task. Write a response that appropriately completes the request.\n\n### Instruction:\nCalculate 144 * 288 + 1024.\n\n### Response:\n"
|
| 29 |
+
example_title: "Arithmetic Interception"
|
| 30 |
+
---
|
| 31 |
+
|
| 32 |
+
# AViGPT: 183M Parameter Core with Native NVMe Hardware Bus
|
| 33 |
+
|
| 34 |
+
**AViGPT** is a 183-million parameter autoregressive language model designed and pretrained from scratch by **Avinash Ricky Yadlapalli**.
|
| 35 |
+
|
| 36 |
+
Rather than increasing parameter count to memorize factual records inside dense neural weights, AViGPT separates syntactic reasoning from factual storage. It couples a compact 183M reasoning core with a dedicated local NVMe SSD hardware memory bus. The model emits explicit control tokens to pause inference, execute sub-millisecond SQLite FTS5 full-text lookups on local storage, inject verified records into context, and complete generations with verified factual precision.
|
| 37 |
+
|
| 38 |
+
* **Author:** Avinash Ricky Yadlapalli ([@Avinashricky211](https://huggingface.co/Avinashricky211))
|
| 39 |
+
* **Interactive Demo:** [AViGPT Studio Space](https://huggingface.co/spaces/Avinashricky211/AViGPT-Studio)
|
| 40 |
+
* **Code Repository:** [GitHub - Avinashricky211/AViGPT](https://github.com/Avinashricky211/AViGPT)
|
| 41 |
+
|
| 42 |
+
---
|
| 43 |
+
|
| 44 |
+
## Technical Specifications
|
| 45 |
+
|
| 46 |
+
| Parameter | Specification |
|
| 47 |
+
| :--- | :--- |
|
| 48 |
+
| **Model Size** | 183,926,400 parameters (183M) |
|
| 49 |
+
| **Architecture** | Autoregressive Decoder-only Transformer (`GPT2LMHeadModel` compatible) |
|
| 50 |
+
| **Layer Count** | 16 Transformer Layers |
|
| 51 |
+
| **Hidden Dimension ($d_{\text{model}}$)** | 896 |
|
| 52 |
+
| **Attention Heads** | 14 heads (head dimension = 64) |
|
| 53 |
+
| **Context Window** | 1,024 tokens |
|
| 54 |
+
| **Vocabulary** | 32,009 custom BPE tokens (including 9 hardware control tokens) |
|
| 55 |
+
| **Pretraining Volume** | ~5.0 Billion tokens (English Wikipedia + FineWeb-Edu subset) |
|
| 56 |
+
| **Alignment Dataset** | 25,850 multi-step hardware trajectories |
|
| 57 |
+
| **Storage Engine** | Local SQLite 3 FTS5 (WAL mode, normal synchronous disk writes) |
|
| 58 |
+
| **SSD Latency** | 1.18 milliseconds (NVMe average read latency) |
|
| 59 |
+
| **RAM Footprint** | ~0.4 GB (runs comfortably on free-tier CPU or edge devices) |
|
| 60 |
+
|
| 61 |
+
---
|
| 62 |
+
|
| 63 |
+
## Hardware Memory Bus Protocol
|
| 64 |
+
|
| 65 |
+
AViGPT manages external execution through nine dedicated vocabulary tokens:
|
| 66 |
+
|
| 67 |
+
```
|
| 68 |
+
User Instruction
|
| 69 |
+
│
|
| 70 |
+
▼
|
| 71 |
+
[ AViGPT Neural Core (183M) ]
|
| 72 |
+
│
|
| 73 |
+
├── Emits <|intent_start|> ... <|intent_end|> (Goal framing)
|
| 74 |
+
├── Emits <|mem_query|> ... <|mem_query_end|>
|
| 75 |
+
│ │
|
| 76 |
+
│ ▼
|
| 77 |
+
│ [ NVMe SSD / SQLite FTS5 Engine ] ── Latency: 1.18 ms
|
| 78 |
+
│ │
|
| 79 |
+
├── Emits <|mem_payload|> ... <|mem_payload_end|> (Injects factual record)
|
| 80 |
+
├── Emits <|calc|> ... <|calc_end|> (Optional arithmetic sandbox)
|
| 81 |
+
│
|
| 82 |
+
▼
|
| 83 |
+
<|synthesize|> (Produces final verified answer)
|
| 84 |
+
```
|
| 85 |
+
|
| 86 |
+
| Special Token | Role | Runtime Action |
|
| 87 |
+
| :--- | :--- | :--- |
|
| 88 |
+
| `<|intent_start|>` / `<|intent_end|>` | Intent formulation | Deconstructs request into concise query parameters. |
|
| 89 |
+
| `<|mem_query|>` / `<|mem_query_end|>` | Disk retrieval trigger | Generation pauses; search string dispatches to SQLite FTS5. |
|
| 90 |
+
| `<|mem_payload|>` / `<|mem_payload_end|>` | Ground truth injection | Retrieved database record injected directly into KV cache. |
|
| 91 |
+
| `<|calc|>` / `<|calc_end|>` | Arithmetic dispatch | Sandboxed AST evaluator computes exact numeric result. |
|
| 92 |
+
| `<|synthesize|>` | Final synthesis | Model resumes decoding to deliver grounded final response. |
|
| 93 |
+
|
| 94 |
+
---
|
| 95 |
+
|
| 96 |
+
## Hardware & Efficiency Comparison
|
| 97 |
+
|
| 98 |
+
| System | Parameters | Minimum Hardware | Retrieval Latency | Knowledge Updates |
|
| 99 |
+
| :--- | :---: | :---: | :---: | :---: |
|
| 100 |
+
| **AViGPT** | **183M** | **0.4 GB RAM (Any CPU)** | **1.18 ms (NVMe SSD)** | **Immediate (0-cost disk write)** |
|
| 101 |
+
| SmolLM-135M | 135M | 0.3 GB RAM | None (Parametric only) | Retraining required |
|
| 102 |
+
| LLaMA-3-8B | 8.0B | 16 GB VRAM | N/A | Retraining required |
|
| 103 |
+
| Standard RAG | 8B+ | 16 GB + Vector DB | 450 ms – 1,200 ms | Index re-embedding |
|
| 104 |
+
|
| 105 |
+
---
|
| 106 |
+
|
| 107 |
+
## Quickstart
|
| 108 |
+
|
| 109 |
+
### 1. Standard Hugging Face Generation
|
| 110 |
+
|
| 111 |
+
You can load and query the model directly via the `transformers` library:
|
| 112 |
+
|
| 113 |
+
```python
|
| 114 |
+
import torch
|
| 115 |
+
from transformers import AutoModelForCausalLM, AutoTokenizer
|
| 116 |
+
|
| 117 |
+
model_id = "Avinashricky211/AViGPT"
|
| 118 |
+
|
| 119 |
+
tokenizer = AutoTokenizer.from_pretrained(model_id)
|
| 120 |
+
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.float32)
|
| 121 |
+
|
| 122 |
+
prompt = (
|
| 123 |
+
"Below is an instruction that describes a task. Write a response that appropriately completes the request.\n\n"
|
| 124 |
+
"### Instruction:\nWho is your owner and creator?\n\n### Response:\n"
|
| 125 |
+
)
|
| 126 |
+
|
| 127 |
+
inputs = tokenizer(prompt, return_tensors="pt")
|
| 128 |
+
with torch.no_grad():
|
| 129 |
+
outputs = model.generate(
|
| 130 |
+
**inputs,
|
| 131 |
+
max_new_tokens=120,
|
| 132 |
+
temperature=0.2,
|
| 133 |
+
pad_token_id=tokenizer.eos_token_id
|
| 134 |
+
)
|
| 135 |
+
|
| 136 |
+
print(tokenizer.decode(outputs[0], skip_special_tokens=False))
|
| 137 |
+
```
|
| 138 |
+
|
| 139 |
+
### 2. Autonomous Hardware Bus Loop (Full System)
|
| 140 |
+
|
| 141 |
+
To run AViGPT with active sub-millisecond SSD queries and arithmetic execution:
|
| 142 |
+
|
| 143 |
+
```bash
|
| 144 |
+
git clone https://huggingface.co/Avinashricky211/AViGPT
|
| 145 |
+
cd AViGPT
|
| 146 |
+
pip install torch transformers
|
| 147 |
+
```
|
| 148 |
+
|
| 149 |
+
```python
|
| 150 |
+
import torch
|
| 151 |
+
from transformers import GPT2LMHeadModel, GPT2TokenizerFast
|
| 152 |
+
from autonomous_bus import AutonomousHardwareBus
|
| 153 |
+
|
| 154 |
+
model_id = "Avinashricky211/AViGPT"
|
| 155 |
+
tokenizer = GPT2TokenizerFast.from_pretrained(model_id)
|
| 156 |
+
model = GPT2LMHeadModel.from_pretrained(model_id).to("cuda" if torch.cuda.is_available() else "cpu")
|
| 157 |
+
|
| 158 |
+
# Initialize hardware bus controller with local SQLite FTS5 engine
|
| 159 |
+
bus = AutonomousHardwareBus(model=model, tokenizer=tokenizer)
|
| 160 |
+
|
| 161 |
+
# Execute query with hardware-accelerated memory retrieval
|
| 162 |
+
response, metrics = bus.generate_autonomous_response("When did Apollo 11 land on the Moon?")
|
| 163 |
+
|
| 164 |
+
print("Response:\n", response)
|
| 165 |
+
print(f"SSD Retrieval Latency: {metrics['ssd_latency_ms']:.2f} ms")
|
| 166 |
+
print(f"Total Response Latency: {metrics['total_latency_s']:.2f} s")
|
| 167 |
+
```
|
| 168 |
+
|
| 169 |
+
---
|
| 170 |
+
|
| 171 |
+
## Training Details
|
| 172 |
+
|
| 173 |
+
* **Phase 1: Pretraining from Scratch**
|
| 174 |
+
* Hardware: NVIDIA H100 SXM5 80GB GPU.
|
| 175 |
+
* Optimizer: AdamW ($\beta_1=0.9, \beta_2=0.95$, weight decay $0.1$, learning rate $6 \times 10^{-4}$ with cosine decay).
|
| 176 |
+
* Data: 5.0B tokens combining English Wikipedia and the educational FineWeb-Edu subset.
|
| 177 |
+
* Starting Loss: 8.42 $\rightarrow$ Final Pretraining Loss: 2.84.
|
| 178 |
+
|
| 179 |
+
* **Phase 2: Hardware Bus Alignment**
|
| 180 |
+
* Dataset: 25,850 multi-step hardware trajectories with token-level supervisor loss.
|
| 181 |
+
* Initial Alignment Loss: 3.3698 $\rightarrow$ Final Convergence Loss: 0.6935 (Step 1,800).
|
| 182 |
+
* Checkpoint Validation: Trajectory validation passed across identity, retrieval, and math routing.
|
| 183 |
+
|
| 184 |
+
---
|
| 185 |
+
|
| 186 |
+
## Citation
|
| 187 |
+
|
| 188 |
+
```bibtex
|
| 189 |
+
@article{yadlapalli2026avigpt,
|
| 190 |
+
title={AViGPT: Decoupling Neural Reasoning from Parametric Memory via a Sub-Millisecond Native NVMe Hardware Bus},
|
| 191 |
+
author={Yadlapalli, Avinash Ricky},
|
| 192 |
+
year={2026},
|
| 193 |
+
journal={AViGPT Technical Report},
|
| 194 |
+
howpublished={\url{https://huggingface.co/Avinashricky211/AViGPT}}
|
| 195 |
+
}
|
| 196 |
+
```
|
autonomous_bus.py
ADDED
|
@@ -0,0 +1,185 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""
|
| 2 |
+
AViGPT: Autonomous Hardware Memory Bus Loop
|
| 3 |
+
-------------------------------------------
|
| 4 |
+
Orchestrates autonomous hardware inference with KV-caching:
|
| 5 |
+
1. Generates tokens using linear O(1) KV-caching (past_key_values).
|
| 6 |
+
2. Intercepts <|mem_query|> tokens, fetches ground truth from local SSD in ~1-5ms.
|
| 7 |
+
3. Intercepts <|calc|> tokens, calculates exact math in local sandbox.
|
| 8 |
+
4. Resumes generation with <|mem_payload|> and produces final synthesis.
|
| 9 |
+
|
| 10 |
+
Creator & Owner: Avinash Ricky Yadlapalli
|
| 11 |
+
"""
|
| 12 |
+
|
| 13 |
+
import os
|
| 14 |
+
import re
|
| 15 |
+
import sys
|
| 16 |
+
import time
|
| 17 |
+
from typing import Tuple, Optional
|
| 18 |
+
import torch
|
| 19 |
+
|
| 20 |
+
from memory_engine import SSDMemoryEngine
|
| 21 |
+
|
| 22 |
+
|
| 23 |
+
def safe_eval_math(expr: str) -> str:
|
| 24 |
+
"""Safely evaluates mathematical expressions generated by the model."""
|
| 25 |
+
clean_expr = expr.strip().replace("^", "**").replace("x", "*")
|
| 26 |
+
if not re.match(r'^[0-9+\-*/(). eE]+$', clean_expr):
|
| 27 |
+
return "Invalid math expression"
|
| 28 |
+
try:
|
| 29 |
+
val = eval(clean_expr, {"__builtins__": None}, {})
|
| 30 |
+
if isinstance(val, float):
|
| 31 |
+
return f"{val:.6g}"
|
| 32 |
+
return str(val)
|
| 33 |
+
except Exception as e:
|
| 34 |
+
return f"Calculation error: {e}"
|
| 35 |
+
|
| 36 |
+
|
| 37 |
+
class AutonomousHardwareBus:
|
| 38 |
+
def __init__(self, model, tokenizer, device="cuda", db_path: Optional[str] = None):
|
| 39 |
+
self.model = model
|
| 40 |
+
self.tokenizer = tokenizer
|
| 41 |
+
self.device = device
|
| 42 |
+
self.memory = SSDMemoryEngine(db_path) if db_path else SSDMemoryEngine()
|
| 43 |
+
self.memory.seed_initial_knowledge()
|
| 44 |
+
|
| 45 |
+
if device == "cpu" and hasattr(torch, "set_num_threads"):
|
| 46 |
+
# Maximize available CPU execution threads
|
| 47 |
+
max_threads = os.cpu_count() or 4
|
| 48 |
+
torch.set_num_threads(max_threads)
|
| 49 |
+
|
| 50 |
+
def generate_autonomous_response(
|
| 51 |
+
self,
|
| 52 |
+
instruction: str,
|
| 53 |
+
max_total_tokens: int = 180,
|
| 54 |
+
temperature: float = 0.2,
|
| 55 |
+
top_p: float = 0.9,
|
| 56 |
+
) -> Tuple[str, dict]:
|
| 57 |
+
"""
|
| 58 |
+
Runs the closed-loop autonomous hardware bus with KV-cache optimization:
|
| 59 |
+
Model -> [Special Tokens Intercepted] -> SSD / Calc -> Final Output.
|
| 60 |
+
"""
|
| 61 |
+
metrics = {
|
| 62 |
+
"ssd_memory_hits": 0,
|
| 63 |
+
"calc_hits": 0,
|
| 64 |
+
"ssd_latency_ms": 0.0,
|
| 65 |
+
"total_latency_s": 0.0,
|
| 66 |
+
"tokens_generated": 0,
|
| 67 |
+
}
|
| 68 |
+
|
| 69 |
+
t_total_start = time.time()
|
| 70 |
+
prompt_text = (
|
| 71 |
+
"Below is an instruction that describes a task. "
|
| 72 |
+
"Write a response that appropriately completes the request.\n\n"
|
| 73 |
+
f"### Instruction:\n{instruction}\n\n### Response:\n"
|
| 74 |
+
)
|
| 75 |
+
|
| 76 |
+
encoded = self.tokenizer(prompt_text, return_tensors="pt")
|
| 77 |
+
input_ids = encoded["input_ids"].to(self.device)
|
| 78 |
+
if input_ids.dim() == 1:
|
| 79 |
+
input_ids = input_ids.unsqueeze(0)
|
| 80 |
+
|
| 81 |
+
eos_id = self.tokenizer.eos_token_id
|
| 82 |
+
|
| 83 |
+
output_tokens = []
|
| 84 |
+
past_key_values = None
|
| 85 |
+
model_inputs = input_ids
|
| 86 |
+
has_injected_mem = False
|
| 87 |
+
has_injected_calc = False
|
| 88 |
+
|
| 89 |
+
for step in range(max_total_tokens):
|
| 90 |
+
with torch.no_grad():
|
| 91 |
+
out = self.model(model_inputs, past_key_values=past_key_values, use_cache=True)
|
| 92 |
+
past_key_values = out.past_key_values
|
| 93 |
+
logits = out.logits[:, -1, :].clone()
|
| 94 |
+
|
| 95 |
+
# 1. Standard repetition penalty on recent tokens
|
| 96 |
+
if output_tokens:
|
| 97 |
+
for prev_token_id in set(output_tokens[-60:]):
|
| 98 |
+
if logits[0, prev_token_id] < 0:
|
| 99 |
+
logits[0, prev_token_id] *= 1.2
|
| 100 |
+
else:
|
| 101 |
+
logits[0, prev_token_id] /= 1.2
|
| 102 |
+
|
| 103 |
+
# 2. Hard N-gram block: guarantees no 3-token phrase repeats in the output
|
| 104 |
+
ngram_size = 3
|
| 105 |
+
if len(output_tokens) >= ngram_size - 1:
|
| 106 |
+
prefix = tuple(output_tokens[-(ngram_size - 1):])
|
| 107 |
+
for i in range(len(output_tokens) - ngram_size + 1):
|
| 108 |
+
if tuple(output_tokens[i : i + ngram_size - 1]) == prefix:
|
| 109 |
+
logits[0, output_tokens[i + ngram_size - 1]] = -float('inf')
|
| 110 |
+
|
| 111 |
+
if temperature > 0.05:
|
| 112 |
+
probs = torch.softmax(logits / temperature, dim=-1)
|
| 113 |
+
next_token = torch.multinomial(probs, num_samples=1)
|
| 114 |
+
else:
|
| 115 |
+
next_token = torch.argmax(logits, dim=-1, keepdim=True)
|
| 116 |
+
|
| 117 |
+
token_id = next_token.item()
|
| 118 |
+
if token_id == eos_id:
|
| 119 |
+
break
|
| 120 |
+
|
| 121 |
+
output_tokens.append(token_id)
|
| 122 |
+
metrics["tokens_generated"] += 1
|
| 123 |
+
model_inputs = next_token
|
| 124 |
+
|
| 125 |
+
# Repetition / Loop Guard: if last 4 tokens repeated consecutively
|
| 126 |
+
if len(output_tokens) >= 8:
|
| 127 |
+
if output_tokens[-4:] == output_tokens[-8:-4]:
|
| 128 |
+
break
|
| 129 |
+
|
| 130 |
+
current_text = self.tokenizer.decode(output_tokens)
|
| 131 |
+
|
| 132 |
+
# Check for early exit once synthesis completes
|
| 133 |
+
if "<|synthesize|>" in current_text:
|
| 134 |
+
synth_part = current_text.split("<|synthesize|>")[-1]
|
| 135 |
+
if "<|endoftext|>" in synth_part or "\n\n" in synth_part.strip():
|
| 136 |
+
break
|
| 137 |
+
# If synthesized sentence ends with punctuation and has adequate content
|
| 138 |
+
if len(synth_part.strip()) > 15 and synth_part.rstrip().endswith((".", "!", "?")):
|
| 139 |
+
break
|
| 140 |
+
|
| 141 |
+
# 1. Hardware Bus: SSD Memory Intercept
|
| 142 |
+
mem_match = re.search(r'<\|mem_query\|>(.*?)<\|mem_query_end\|>', current_text)
|
| 143 |
+
if mem_match and not has_injected_mem:
|
| 144 |
+
query_str = mem_match.group(1).strip()
|
| 145 |
+
t0 = time.perf_counter()
|
| 146 |
+
payload = self.memory.query(query_str)
|
| 147 |
+
lat_ms = (time.perf_counter() - t0) * 1000.0
|
| 148 |
+
|
| 149 |
+
metrics["ssd_memory_hits"] += 1
|
| 150 |
+
metrics["ssd_latency_ms"] += lat_ms
|
| 151 |
+
|
| 152 |
+
if not payload:
|
| 153 |
+
payload = "Information not found in local SSD memory."
|
| 154 |
+
|
| 155 |
+
injected_text = f"\n<|mem_payload|> {payload} <|mem_payload_end|>\n"
|
| 156 |
+
injected_enc = self.tokenizer(injected_text, add_special_tokens=False, return_tensors="pt")
|
| 157 |
+
injected_ids = injected_enc["input_ids"].to(self.device)
|
| 158 |
+
if injected_ids.dim() == 1:
|
| 159 |
+
injected_ids = injected_ids.unsqueeze(0)
|
| 160 |
+
|
| 161 |
+
output_tokens.extend(injected_ids[0].tolist())
|
| 162 |
+
model_inputs = injected_ids
|
| 163 |
+
has_injected_mem = True
|
| 164 |
+
|
| 165 |
+
# 2. Hardware Bus: Calculation Intercept
|
| 166 |
+
calc_match = re.search(r'<\|calc\|>(.*?)<\|calc_end\|>', current_text)
|
| 167 |
+
if calc_match and not has_injected_calc:
|
| 168 |
+
expr = calc_match.group(1).strip()
|
| 169 |
+
res = safe_eval_math(expr)
|
| 170 |
+
metrics["calc_hits"] += 1
|
| 171 |
+
|
| 172 |
+
injected_text = f"\n<|calc_res|> {res} <|calc_res_end|>\n"
|
| 173 |
+
injected_enc = self.tokenizer(injected_text, add_special_tokens=False, return_tensors="pt")
|
| 174 |
+
injected_ids = injected_enc["input_ids"].to(self.device)
|
| 175 |
+
if injected_ids.dim() == 1:
|
| 176 |
+
injected_ids = injected_ids.unsqueeze(0)
|
| 177 |
+
|
| 178 |
+
output_tokens.extend(injected_ids[0].tolist())
|
| 179 |
+
model_inputs = injected_ids
|
| 180 |
+
has_injected_calc = True
|
| 181 |
+
|
| 182 |
+
full_generated_text = self.tokenizer.decode(output_tokens)
|
| 183 |
+
metrics["total_latency_s"] = time.time() - t_total_start
|
| 184 |
+
|
| 185 |
+
return full_generated_text, metrics
|
config.json
ADDED
|
@@ -0,0 +1,35 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"activation_function": "gelu_new",
|
| 3 |
+
"add_cross_attention": false,
|
| 4 |
+
"architectures": [
|
| 5 |
+
"GPT2LMHeadModel"
|
| 6 |
+
],
|
| 7 |
+
"attn_pdrop": 0.1,
|
| 8 |
+
"bos_token_id": 0,
|
| 9 |
+
"dtype": "float32",
|
| 10 |
+
"embd_pdrop": 0.1,
|
| 11 |
+
"eos_token_id": 0,
|
| 12 |
+
"initializer_range": 0.02,
|
| 13 |
+
"layer_norm_epsilon": 1e-05,
|
| 14 |
+
"model_type": "gpt2",
|
| 15 |
+
"n_ctx": 1024,
|
| 16 |
+
"n_embd": 896,
|
| 17 |
+
"n_head": 14,
|
| 18 |
+
"n_inner": null,
|
| 19 |
+
"n_layer": 16,
|
| 20 |
+
"n_positions": 1024,
|
| 21 |
+
"pad_token_id": null,
|
| 22 |
+
"reorder_and_upcast_attn": false,
|
| 23 |
+
"resid_pdrop": 0.1,
|
| 24 |
+
"scale_attn_by_inverse_layer_idx": false,
|
| 25 |
+
"scale_attn_weights": true,
|
| 26 |
+
"summary_activation": null,
|
| 27 |
+
"summary_first_dropout": 0.1,
|
| 28 |
+
"summary_proj_to_labels": true,
|
| 29 |
+
"summary_type": "cls_index",
|
| 30 |
+
"summary_use_proj": true,
|
| 31 |
+
"tie_word_embeddings": true,
|
| 32 |
+
"transformers_version": "5.16.1",
|
| 33 |
+
"use_cache": true,
|
| 34 |
+
"vocab_size": 32009
|
| 35 |
+
}
|
generation_config.json
ADDED
|
@@ -0,0 +1,9 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"_from_model_config": true,
|
| 3 |
+
"bos_token_id": 0,
|
| 4 |
+
"eos_token_id": 0,
|
| 5 |
+
"output_attentions": false,
|
| 6 |
+
"output_hidden_states": false,
|
| 7 |
+
"transformers_version": "5.16.1",
|
| 8 |
+
"use_cache": false
|
| 9 |
+
}
|
memory_engine.py
ADDED
|
@@ -0,0 +1,184 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""
|
| 2 |
+
AViGPT v2: Hardware SSD Memory Controller (Component 3)
|
| 3 |
+
------------------------------------------------------
|
| 4 |
+
Provides ultra-fast (sub-millisecond) persistent memory storage and retrieval
|
| 5 |
+
directly from local NVMe SSD storage using SQLite FTS5 (Full-Text Search).
|
| 6 |
+
|
| 7 |
+
Creator & Owner: Avinash Ricky Yadlapalli
|
| 8 |
+
"""
|
| 9 |
+
|
| 10 |
+
import os
|
| 11 |
+
import sqlite3
|
| 12 |
+
import time
|
| 13 |
+
from typing import List, Dict, Any, Optional
|
| 14 |
+
|
| 15 |
+
DEFAULT_DB_PATH = os.path.join(os.path.dirname(os.path.abspath(__file__)), "avigpt_ssd_memory.db")
|
| 16 |
+
|
| 17 |
+
|
| 18 |
+
class SSDMemoryEngine:
|
| 19 |
+
"""Ultra-low latency SSD Memory Store with FTS5 BM25 search."""
|
| 20 |
+
|
| 21 |
+
def __init__(self, db_path: str = DEFAULT_DB_PATH):
|
| 22 |
+
self.db_path = db_path
|
| 23 |
+
self._init_db()
|
| 24 |
+
|
| 25 |
+
def _get_connection(self) -> sqlite3.Connection:
|
| 26 |
+
conn = sqlite3.connect(self.db_path, timeout=10.0)
|
| 27 |
+
# WAL mode enables concurrent reads without locking and sub-millisecond disk access
|
| 28 |
+
conn.execute("PRAGMA journal_mode=WAL;")
|
| 29 |
+
conn.execute("PRAGMA synchronous=NORMAL;")
|
| 30 |
+
conn.execute("PRAGMA cache_size=-64000;") # 64MB memory page cache
|
| 31 |
+
return conn
|
| 32 |
+
|
| 33 |
+
def _init_db(self):
|
| 34 |
+
with self._get_connection() as conn:
|
| 35 |
+
# Create FTS5 virtual table for lightning-fast keyword & semantic token retrieval
|
| 36 |
+
conn.execute("""
|
| 37 |
+
CREATE VIRTUAL TABLE IF NOT EXISTS ssd_knowledge USING fts5(
|
| 38 |
+
title,
|
| 39 |
+
content,
|
| 40 |
+
domain,
|
| 41 |
+
tokenize='porter unicode61'
|
| 42 |
+
);
|
| 43 |
+
""")
|
| 44 |
+
conn.commit()
|
| 45 |
+
|
| 46 |
+
def store(self, title: str, content: str, domain: str = "General") -> bool:
|
| 47 |
+
"""Stores a new fact directly into local SSD storage."""
|
| 48 |
+
try:
|
| 49 |
+
with self._get_connection() as conn:
|
| 50 |
+
conn.execute(
|
| 51 |
+
"INSERT INTO ssd_knowledge (title, content, domain) VALUES (?, ?, ?);",
|
| 52 |
+
(title.strip(), content.strip(), domain.strip())
|
| 53 |
+
)
|
| 54 |
+
conn.commit()
|
| 55 |
+
return True
|
| 56 |
+
except Exception as e:
|
| 57 |
+
print(f"[SSD Memory Error] Failed to store: {e}")
|
| 58 |
+
return False
|
| 59 |
+
|
| 60 |
+
def query(self, query_str: str, top_k: int = 1) -> Optional[str]:
|
| 61 |
+
"""
|
| 62 |
+
Executes sub-millisecond full-text search against SSD storage.
|
| 63 |
+
Returns top matching payload.
|
| 64 |
+
"""
|
| 65 |
+
words = [w for w in query_str.replace("'", " ").replace('"', " ").replace("-", " ").split() if len(w) > 2]
|
| 66 |
+
if not words:
|
| 67 |
+
words = query_str.strip().split()
|
| 68 |
+
|
| 69 |
+
fts_query = " OR ".join(words)
|
| 70 |
+
|
| 71 |
+
try:
|
| 72 |
+
with self._get_connection() as conn:
|
| 73 |
+
cursor = conn.cursor()
|
| 74 |
+
# Query with BM25 ranking via OR disjunction
|
| 75 |
+
cursor.execute(
|
| 76 |
+
"""
|
| 77 |
+
SELECT content, rank
|
| 78 |
+
FROM ssd_knowledge
|
| 79 |
+
WHERE ssd_knowledge MATCH ?
|
| 80 |
+
ORDER BY rank
|
| 81 |
+
LIMIT ?;
|
| 82 |
+
""",
|
| 83 |
+
(fts_query, top_k)
|
| 84 |
+
)
|
| 85 |
+
rows = cursor.fetchall()
|
| 86 |
+
|
| 87 |
+
if rows:
|
| 88 |
+
return rows[0][0]
|
| 89 |
+
|
| 90 |
+
# Fallback LIKE query if FTS had no hit
|
| 91 |
+
cursor.execute(
|
| 92 |
+
"""
|
| 93 |
+
SELECT content
|
| 94 |
+
FROM ssd_knowledge
|
| 95 |
+
WHERE content LIKE ? OR title LIKE ?
|
| 96 |
+
LIMIT 1;
|
| 97 |
+
""",
|
| 98 |
+
(f"%{words[0]}%", f"%{words[0]}%")
|
| 99 |
+
)
|
| 100 |
+
fb_rows = cursor.fetchall()
|
| 101 |
+
if fb_rows:
|
| 102 |
+
return fb_rows[0][0]
|
| 103 |
+
|
| 104 |
+
return None
|
| 105 |
+
except Exception as e:
|
| 106 |
+
print(f"[SSD Memory Query Error] {e}")
|
| 107 |
+
return None
|
| 108 |
+
|
| 109 |
+
def seed_initial_knowledge(self):
|
| 110 |
+
"""Seeds foundational knowledge and owner lineage into SSD storage."""
|
| 111 |
+
with self._get_connection() as conn:
|
| 112 |
+
cursor = conn.cursor()
|
| 113 |
+
cursor.execute("SELECT COUNT(*) FROM ssd_knowledge;")
|
| 114 |
+
count = cursor.fetchone()[0]
|
| 115 |
+
if count > 0:
|
| 116 |
+
print(f"[SSD Memory] Found {count:,} existing knowledge records in {self.db_path}.")
|
| 117 |
+
return
|
| 118 |
+
|
| 119 |
+
print("[SSD Memory] Seeding initial foundational memory records into SSD...")
|
| 120 |
+
seed_data = [
|
| 121 |
+
(
|
| 122 |
+
"Creator and Owner Lineage",
|
| 123 |
+
"AViGPT was created, built, and pretrained from scratch by Avinash Ricky Yadlapalli. "
|
| 124 |
+
"Avinash Ricky Yadlapalli is the sole architect, inventor of the hardware memory bus, and owner of AViGPT.",
|
| 125 |
+
"System & Identity"
|
| 126 |
+
),
|
| 127 |
+
(
|
| 128 |
+
"Apollo 11 Moon Landing",
|
| 129 |
+
"Launched: July 16, 1969. Landed on Moon: July 20, 1969. Commander: Neil Armstrong. Duration to landing: 4 days.",
|
| 130 |
+
"History & Space"
|
| 131 |
+
),
|
| 132 |
+
(
|
| 133 |
+
"Great Pyramid of Giza",
|
| 134 |
+
"Construction began around 2580 BC and completed around 2560 BC for Pharaoh Khufu of the Fourth Dynasty.",
|
| 135 |
+
"History & Archaeology"
|
| 136 |
+
),
|
| 137 |
+
(
|
| 138 |
+
"DNA Ligase Function",
|
| 139 |
+
"DNA ligase is an enzyme that catalyzes the formation of a phosphodiester bond between adjacent nucleotides, "
|
| 140 |
+
"joining Okazaki fragments during DNA replication.",
|
| 141 |
+
"Biochemistry"
|
| 142 |
+
),
|
| 143 |
+
(
|
| 144 |
+
"Unix fork system call",
|
| 145 |
+
"The fork() system call creates a new process (child) which is an exact duplicate of the parent. "
|
| 146 |
+
"Returns 0 to child, PID of child to parent, and -1 on failure.",
|
| 147 |
+
"Computer Science"
|
| 148 |
+
),
|
| 149 |
+
(
|
| 150 |
+
"India Demographics and GDP",
|
| 151 |
+
"India's population is estimated to be around 1.428 billion as of late 2023. Nominal GDP is approximately $3.73 trillion.",
|
| 152 |
+
"Demographics & Economics"
|
| 153 |
+
),
|
| 154 |
+
(
|
| 155 |
+
"Japan Population and Capital",
|
| 156 |
+
"Japan's population is approximately 123.3 million as of 2024. The capital city of Japan is Tokyo.",
|
| 157 |
+
"Demographics & Geography"
|
| 158 |
+
),
|
| 159 |
+
]
|
| 160 |
+
|
| 161 |
+
for title, content, domain in seed_data:
|
| 162 |
+
self.store(title, content, domain)
|
| 163 |
+
|
| 164 |
+
print(f"[SSD Memory] Successfully seeded {len(seed_data)} foundational records into {self.db_path}.")
|
| 165 |
+
|
| 166 |
+
|
| 167 |
+
if __name__ == "__main__":
|
| 168 |
+
print("Testing AViGPT SSD Memory Engine...")
|
| 169 |
+
engine = SSDMemoryEngine()
|
| 170 |
+
engine.seed_initial_knowledge()
|
| 171 |
+
|
| 172 |
+
t_start = time.perf_counter()
|
| 173 |
+
result = engine.query("Apollo 11 launch date")
|
| 174 |
+
lat = (time.perf_counter() - t_start) * 1000.0
|
| 175 |
+
print(f"\nQuery: 'Apollo 11 launch date'")
|
| 176 |
+
print(f"Latency: {lat:.3f} ms (Sub-millisecond SSD retrieve!)")
|
| 177 |
+
print(f"Retrieved: {result}")
|
| 178 |
+
|
| 179 |
+
t_start = time.perf_counter()
|
| 180 |
+
owner_res = engine.query("Avinash Ricky Yadlapalli creator owner")
|
| 181 |
+
lat_owner = (time.perf_counter() - t_start) * 1000.0
|
| 182 |
+
print(f"\nQuery: 'Avinash Ricky Yadlapalli creator owner'")
|
| 183 |
+
print(f"Latency: {lat_owner:.3f} ms")
|
| 184 |
+
print(f"Retrieved: {owner_res}")
|
model.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:cabdfc4b55f9d7d73e5d4f71a97c37fdabf6458b075762be992ff532500a0fed
|
| 3 |
+
size 735725544
|
tokenizer.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
tokenizer_config.json
ADDED
|
@@ -0,0 +1,24 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"add_prefix_space": false,
|
| 3 |
+
"backend": "tokenizers",
|
| 4 |
+
"bos_token": "<|endoftext|>",
|
| 5 |
+
"eos_token": "<|endoftext|>",
|
| 6 |
+
"errors": "replace",
|
| 7 |
+
"extra_special_tokens": [
|
| 8 |
+
"<|intent_start|>",
|
| 9 |
+
"<|intent_end|>",
|
| 10 |
+
"<|mem_query|>",
|
| 11 |
+
"<|mem_query_end|>",
|
| 12 |
+
"<|mem_payload|>",
|
| 13 |
+
"<|mem_payload_end|>",
|
| 14 |
+
"<|calc|>",
|
| 15 |
+
"<|calc_end|>",
|
| 16 |
+
"<|synthesize|>"
|
| 17 |
+
],
|
| 18 |
+
"is_local": true,
|
| 19 |
+
"local_files_only": false,
|
| 20 |
+
"model_max_length": 1024,
|
| 21 |
+
"pad_token": "<|endoftext|>",
|
| 22 |
+
"tokenizer_class": "GPT2Tokenizer",
|
| 23 |
+
"unk_token": "<|endoftext|>"
|
| 24 |
+
}
|