AvinashRicky commited on
Commit
42a0779
·
verified ·
1 Parent(s): 0311ff4

Release AViGPT 183M model weights, config, and Model Card

Browse files
README.md ADDED
@@ -0,0 +1,196 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - en
4
+ license: mit
5
+ library_name: transformers
6
+ tags:
7
+ - text-generation
8
+ - hardware-bus
9
+ - memory-augmented
10
+ - toolformer
11
+ - slm
12
+ - autonomous-agent
13
+ - ssd-memory
14
+ - pytorch
15
+ - gpt2
16
+ - edge-ai
17
+ pipeline_tag: text-generation
18
+ inference:
19
+ parameters:
20
+ temperature: 0.2
21
+ max_new_tokens: 150
22
+ repetition_penalty: 1.15
23
+ widget:
24
+ - text: "Below is an instruction that describes a task. Write a response that appropriately completes the request.\n\n### Instruction:\nWho is your owner and creator?\n\n### Response:\n"
25
+ example_title: "Owner Lineage"
26
+ - text: "Below is an instruction that describes a task. Write a response that appropriately completes the request.\n\n### Instruction:\nWhen was the Apollo 11 Moon landing?\n\n### Response:\n"
27
+ example_title: "SSD Memory Retrieval"
28
+ - text: "Below is an instruction that describes a task. Write a response that appropriately completes the request.\n\n### Instruction:\nCalculate 144 * 288 + 1024.\n\n### Response:\n"
29
+ example_title: "Arithmetic Interception"
30
+ ---
31
+
32
+ # AViGPT: 183M Parameter Core with Native NVMe Hardware Bus
33
+
34
+ **AViGPT** is a 183-million parameter autoregressive language model designed and pretrained from scratch by **Avinash Ricky Yadlapalli**.
35
+
36
+ Rather than increasing parameter count to memorize factual records inside dense neural weights, AViGPT separates syntactic reasoning from factual storage. It couples a compact 183M reasoning core with a dedicated local NVMe SSD hardware memory bus. The model emits explicit control tokens to pause inference, execute sub-millisecond SQLite FTS5 full-text lookups on local storage, inject verified records into context, and complete generations with verified factual precision.
37
+
38
+ * **Author:** Avinash Ricky Yadlapalli ([@Avinashricky211](https://huggingface.co/Avinashricky211))
39
+ * **Interactive Demo:** [AViGPT Studio Space](https://huggingface.co/spaces/Avinashricky211/AViGPT-Studio)
40
+ * **Code Repository:** [GitHub - Avinashricky211/AViGPT](https://github.com/Avinashricky211/AViGPT)
41
+
42
+ ---
43
+
44
+ ## Technical Specifications
45
+
46
+ | Parameter | Specification |
47
+ | :--- | :--- |
48
+ | **Model Size** | 183,926,400 parameters (183M) |
49
+ | **Architecture** | Autoregressive Decoder-only Transformer (`GPT2LMHeadModel` compatible) |
50
+ | **Layer Count** | 16 Transformer Layers |
51
+ | **Hidden Dimension ($d_{\text{model}}$)** | 896 |
52
+ | **Attention Heads** | 14 heads (head dimension = 64) |
53
+ | **Context Window** | 1,024 tokens |
54
+ | **Vocabulary** | 32,009 custom BPE tokens (including 9 hardware control tokens) |
55
+ | **Pretraining Volume** | ~5.0 Billion tokens (English Wikipedia + FineWeb-Edu subset) |
56
+ | **Alignment Dataset** | 25,850 multi-step hardware trajectories |
57
+ | **Storage Engine** | Local SQLite 3 FTS5 (WAL mode, normal synchronous disk writes) |
58
+ | **SSD Latency** | 1.18 milliseconds (NVMe average read latency) |
59
+ | **RAM Footprint** | ~0.4 GB (runs comfortably on free-tier CPU or edge devices) |
60
+
61
+ ---
62
+
63
+ ## Hardware Memory Bus Protocol
64
+
65
+ AViGPT manages external execution through nine dedicated vocabulary tokens:
66
+
67
+ ```
68
+ User Instruction
69
+ │
70
+ ▼
71
+ [ AViGPT Neural Core (183M) ]
72
+ │
73
+ ├── Emits <|intent_start|> ... <|intent_end|> (Goal framing)
74
+ ├── Emits <|mem_query|> ... <|mem_query_end|>
75
+ │ │
76
+ │ ▼
77
+ │ [ NVMe SSD / SQLite FTS5 Engine ] ── Latency: 1.18 ms
78
+ │ │
79
+ ├── Emits <|mem_payload|> ... <|mem_payload_end|> (Injects factual record)
80
+ ├── Emits <|calc|> ... <|calc_end|> (Optional arithmetic sandbox)
81
+ │
82
+ ▼
83
+ <|synthesize|> (Produces final verified answer)
84
+ ```
85
+
86
+ | Special Token | Role | Runtime Action |
87
+ | :--- | :--- | :--- |
88
+ | `<|intent_start|>` / `<|intent_end|>` | Intent formulation | Deconstructs request into concise query parameters. |
89
+ | `<|mem_query|>` / `<|mem_query_end|>` | Disk retrieval trigger | Generation pauses; search string dispatches to SQLite FTS5. |
90
+ | `<|mem_payload|>` / `<|mem_payload_end|>` | Ground truth injection | Retrieved database record injected directly into KV cache. |
91
+ | `<|calc|>` / `<|calc_end|>` | Arithmetic dispatch | Sandboxed AST evaluator computes exact numeric result. |
92
+ | `<|synthesize|>` | Final synthesis | Model resumes decoding to deliver grounded final response. |
93
+
94
+ ---
95
+
96
+ ## Hardware & Efficiency Comparison
97
+
98
+ | System | Parameters | Minimum Hardware | Retrieval Latency | Knowledge Updates |
99
+ | :--- | :---: | :---: | :---: | :---: |
100
+ | **AViGPT** | **183M** | **0.4 GB RAM (Any CPU)** | **1.18 ms (NVMe SSD)** | **Immediate (0-cost disk write)** |
101
+ | SmolLM-135M | 135M | 0.3 GB RAM | None (Parametric only) | Retraining required |
102
+ | LLaMA-3-8B | 8.0B | 16 GB VRAM | N/A | Retraining required |
103
+ | Standard RAG | 8B+ | 16 GB + Vector DB | 450 ms – 1,200 ms | Index re-embedding |
104
+
105
+ ---
106
+
107
+ ## Quickstart
108
+
109
+ ### 1. Standard Hugging Face Generation
110
+
111
+ You can load and query the model directly via the `transformers` library:
112
+
113
+ ```python
114
+ import torch
115
+ from transformers import AutoModelForCausalLM, AutoTokenizer
116
+
117
+ model_id = "Avinashricky211/AViGPT"
118
+
119
+ tokenizer = AutoTokenizer.from_pretrained(model_id)
120
+ model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.float32)
121
+
122
+ prompt = (
123
+ "Below is an instruction that describes a task. Write a response that appropriately completes the request.\n\n"
124
+ "### Instruction:\nWho is your owner and creator?\n\n### Response:\n"
125
+ )
126
+
127
+ inputs = tokenizer(prompt, return_tensors="pt")
128
+ with torch.no_grad():
129
+ outputs = model.generate(
130
+ **inputs,
131
+ max_new_tokens=120,
132
+ temperature=0.2,
133
+ pad_token_id=tokenizer.eos_token_id
134
+ )
135
+
136
+ print(tokenizer.decode(outputs[0], skip_special_tokens=False))
137
+ ```
138
+
139
+ ### 2. Autonomous Hardware Bus Loop (Full System)
140
+
141
+ To run AViGPT with active sub-millisecond SSD queries and arithmetic execution:
142
+
143
+ ```bash
144
+ git clone https://huggingface.co/Avinashricky211/AViGPT
145
+ cd AViGPT
146
+ pip install torch transformers
147
+ ```
148
+
149
+ ```python
150
+ import torch
151
+ from transformers import GPT2LMHeadModel, GPT2TokenizerFast
152
+ from autonomous_bus import AutonomousHardwareBus
153
+
154
+ model_id = "Avinashricky211/AViGPT"
155
+ tokenizer = GPT2TokenizerFast.from_pretrained(model_id)
156
+ model = GPT2LMHeadModel.from_pretrained(model_id).to("cuda" if torch.cuda.is_available() else "cpu")
157
+
158
+ # Initialize hardware bus controller with local SQLite FTS5 engine
159
+ bus = AutonomousHardwareBus(model=model, tokenizer=tokenizer)
160
+
161
+ # Execute query with hardware-accelerated memory retrieval
162
+ response, metrics = bus.generate_autonomous_response("When did Apollo 11 land on the Moon?")
163
+
164
+ print("Response:\n", response)
165
+ print(f"SSD Retrieval Latency: {metrics['ssd_latency_ms']:.2f} ms")
166
+ print(f"Total Response Latency: {metrics['total_latency_s']:.2f} s")
167
+ ```
168
+
169
+ ---
170
+
171
+ ## Training Details
172
+
173
+ * **Phase 1: Pretraining from Scratch**
174
+ * Hardware: NVIDIA H100 SXM5 80GB GPU.
175
+ * Optimizer: AdamW ($\beta_1=0.9, \beta_2=0.95$, weight decay $0.1$, learning rate $6 \times 10^{-4}$ with cosine decay).
176
+ * Data: 5.0B tokens combining English Wikipedia and the educational FineWeb-Edu subset.
177
+ * Starting Loss: 8.42 $\rightarrow$ Final Pretraining Loss: 2.84.
178
+
179
+ * **Phase 2: Hardware Bus Alignment**
180
+ * Dataset: 25,850 multi-step hardware trajectories with token-level supervisor loss.
181
+ * Initial Alignment Loss: 3.3698 $\rightarrow$ Final Convergence Loss: 0.6935 (Step 1,800).
182
+ * Checkpoint Validation: Trajectory validation passed across identity, retrieval, and math routing.
183
+
184
+ ---
185
+
186
+ ## Citation
187
+
188
+ ```bibtex
189
+ @article{yadlapalli2026avigpt,
190
+ title={AViGPT: Decoupling Neural Reasoning from Parametric Memory via a Sub-Millisecond Native NVMe Hardware Bus},
191
+ author={Yadlapalli, Avinash Ricky},
192
+ year={2026},
193
+ journal={AViGPT Technical Report},
194
+ howpublished={\url{https://huggingface.co/Avinashricky211/AViGPT}}
195
+ }
196
+ ```
autonomous_bus.py ADDED
@@ -0,0 +1,185 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """
2
+ AViGPT: Autonomous Hardware Memory Bus Loop
3
+ -------------------------------------------
4
+ Orchestrates autonomous hardware inference with KV-caching:
5
+ 1. Generates tokens using linear O(1) KV-caching (past_key_values).
6
+ 2. Intercepts <|mem_query|> tokens, fetches ground truth from local SSD in ~1-5ms.
7
+ 3. Intercepts <|calc|> tokens, calculates exact math in local sandbox.
8
+ 4. Resumes generation with <|mem_payload|> and produces final synthesis.
9
+
10
+ Creator & Owner: Avinash Ricky Yadlapalli
11
+ """
12
+
13
+ import os
14
+ import re
15
+ import sys
16
+ import time
17
+ from typing import Tuple, Optional
18
+ import torch
19
+
20
+ from memory_engine import SSDMemoryEngine
21
+
22
+
23
+ def safe_eval_math(expr: str) -> str:
24
+ """Safely evaluates mathematical expressions generated by the model."""
25
+ clean_expr = expr.strip().replace("^", "**").replace("x", "*")
26
+ if not re.match(r'^[0-9+\-*/(). eE]+$', clean_expr):
27
+ return "Invalid math expression"
28
+ try:
29
+ val = eval(clean_expr, {"__builtins__": None}, {})
30
+ if isinstance(val, float):
31
+ return f"{val:.6g}"
32
+ return str(val)
33
+ except Exception as e:
34
+ return f"Calculation error: {e}"
35
+
36
+
37
+ class AutonomousHardwareBus:
38
+ def __init__(self, model, tokenizer, device="cuda", db_path: Optional[str] = None):
39
+ self.model = model
40
+ self.tokenizer = tokenizer
41
+ self.device = device
42
+ self.memory = SSDMemoryEngine(db_path) if db_path else SSDMemoryEngine()
43
+ self.memory.seed_initial_knowledge()
44
+
45
+ if device == "cpu" and hasattr(torch, "set_num_threads"):
46
+ # Maximize available CPU execution threads
47
+ max_threads = os.cpu_count() or 4
48
+ torch.set_num_threads(max_threads)
49
+
50
+ def generate_autonomous_response(
51
+ self,
52
+ instruction: str,
53
+ max_total_tokens: int = 180,
54
+ temperature: float = 0.2,
55
+ top_p: float = 0.9,
56
+ ) -> Tuple[str, dict]:
57
+ """
58
+ Runs the closed-loop autonomous hardware bus with KV-cache optimization:
59
+ Model -> [Special Tokens Intercepted] -> SSD / Calc -> Final Output.
60
+ """
61
+ metrics = {
62
+ "ssd_memory_hits": 0,
63
+ "calc_hits": 0,
64
+ "ssd_latency_ms": 0.0,
65
+ "total_latency_s": 0.0,
66
+ "tokens_generated": 0,
67
+ }
68
+
69
+ t_total_start = time.time()
70
+ prompt_text = (
71
+ "Below is an instruction that describes a task. "
72
+ "Write a response that appropriately completes the request.\n\n"
73
+ f"### Instruction:\n{instruction}\n\n### Response:\n"
74
+ )
75
+
76
+ encoded = self.tokenizer(prompt_text, return_tensors="pt")
77
+ input_ids = encoded["input_ids"].to(self.device)
78
+ if input_ids.dim() == 1:
79
+ input_ids = input_ids.unsqueeze(0)
80
+
81
+ eos_id = self.tokenizer.eos_token_id
82
+
83
+ output_tokens = []
84
+ past_key_values = None
85
+ model_inputs = input_ids
86
+ has_injected_mem = False
87
+ has_injected_calc = False
88
+
89
+ for step in range(max_total_tokens):
90
+ with torch.no_grad():
91
+ out = self.model(model_inputs, past_key_values=past_key_values, use_cache=True)
92
+ past_key_values = out.past_key_values
93
+ logits = out.logits[:, -1, :].clone()
94
+
95
+ # 1. Standard repetition penalty on recent tokens
96
+ if output_tokens:
97
+ for prev_token_id in set(output_tokens[-60:]):
98
+ if logits[0, prev_token_id] < 0:
99
+ logits[0, prev_token_id] *= 1.2
100
+ else:
101
+ logits[0, prev_token_id] /= 1.2
102
+
103
+ # 2. Hard N-gram block: guarantees no 3-token phrase repeats in the output
104
+ ngram_size = 3
105
+ if len(output_tokens) >= ngram_size - 1:
106
+ prefix = tuple(output_tokens[-(ngram_size - 1):])
107
+ for i in range(len(output_tokens) - ngram_size + 1):
108
+ if tuple(output_tokens[i : i + ngram_size - 1]) == prefix:
109
+ logits[0, output_tokens[i + ngram_size - 1]] = -float('inf')
110
+
111
+ if temperature > 0.05:
112
+ probs = torch.softmax(logits / temperature, dim=-1)
113
+ next_token = torch.multinomial(probs, num_samples=1)
114
+ else:
115
+ next_token = torch.argmax(logits, dim=-1, keepdim=True)
116
+
117
+ token_id = next_token.item()
118
+ if token_id == eos_id:
119
+ break
120
+
121
+ output_tokens.append(token_id)
122
+ metrics["tokens_generated"] += 1
123
+ model_inputs = next_token
124
+
125
+ # Repetition / Loop Guard: if last 4 tokens repeated consecutively
126
+ if len(output_tokens) >= 8:
127
+ if output_tokens[-4:] == output_tokens[-8:-4]:
128
+ break
129
+
130
+ current_text = self.tokenizer.decode(output_tokens)
131
+
132
+ # Check for early exit once synthesis completes
133
+ if "<|synthesize|>" in current_text:
134
+ synth_part = current_text.split("<|synthesize|>")[-1]
135
+ if "<|endoftext|>" in synth_part or "\n\n" in synth_part.strip():
136
+ break
137
+ # If synthesized sentence ends with punctuation and has adequate content
138
+ if len(synth_part.strip()) > 15 and synth_part.rstrip().endswith((".", "!", "?")):
139
+ break
140
+
141
+ # 1. Hardware Bus: SSD Memory Intercept
142
+ mem_match = re.search(r'<\|mem_query\|>(.*?)<\|mem_query_end\|>', current_text)
143
+ if mem_match and not has_injected_mem:
144
+ query_str = mem_match.group(1).strip()
145
+ t0 = time.perf_counter()
146
+ payload = self.memory.query(query_str)
147
+ lat_ms = (time.perf_counter() - t0) * 1000.0
148
+
149
+ metrics["ssd_memory_hits"] += 1
150
+ metrics["ssd_latency_ms"] += lat_ms
151
+
152
+ if not payload:
153
+ payload = "Information not found in local SSD memory."
154
+
155
+ injected_text = f"\n<|mem_payload|> {payload} <|mem_payload_end|>\n"
156
+ injected_enc = self.tokenizer(injected_text, add_special_tokens=False, return_tensors="pt")
157
+ injected_ids = injected_enc["input_ids"].to(self.device)
158
+ if injected_ids.dim() == 1:
159
+ injected_ids = injected_ids.unsqueeze(0)
160
+
161
+ output_tokens.extend(injected_ids[0].tolist())
162
+ model_inputs = injected_ids
163
+ has_injected_mem = True
164
+
165
+ # 2. Hardware Bus: Calculation Intercept
166
+ calc_match = re.search(r'<\|calc\|>(.*?)<\|calc_end\|>', current_text)
167
+ if calc_match and not has_injected_calc:
168
+ expr = calc_match.group(1).strip()
169
+ res = safe_eval_math(expr)
170
+ metrics["calc_hits"] += 1
171
+
172
+ injected_text = f"\n<|calc_res|> {res} <|calc_res_end|>\n"
173
+ injected_enc = self.tokenizer(injected_text, add_special_tokens=False, return_tensors="pt")
174
+ injected_ids = injected_enc["input_ids"].to(self.device)
175
+ if injected_ids.dim() == 1:
176
+ injected_ids = injected_ids.unsqueeze(0)
177
+
178
+ output_tokens.extend(injected_ids[0].tolist())
179
+ model_inputs = injected_ids
180
+ has_injected_calc = True
181
+
182
+ full_generated_text = self.tokenizer.decode(output_tokens)
183
+ metrics["total_latency_s"] = time.time() - t_total_start
184
+
185
+ return full_generated_text, metrics
config.json ADDED
@@ -0,0 +1,35 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "activation_function": "gelu_new",
3
+ "add_cross_attention": false,
4
+ "architectures": [
5
+ "GPT2LMHeadModel"
6
+ ],
7
+ "attn_pdrop": 0.1,
8
+ "bos_token_id": 0,
9
+ "dtype": "float32",
10
+ "embd_pdrop": 0.1,
11
+ "eos_token_id": 0,
12
+ "initializer_range": 0.02,
13
+ "layer_norm_epsilon": 1e-05,
14
+ "model_type": "gpt2",
15
+ "n_ctx": 1024,
16
+ "n_embd": 896,
17
+ "n_head": 14,
18
+ "n_inner": null,
19
+ "n_layer": 16,
20
+ "n_positions": 1024,
21
+ "pad_token_id": null,
22
+ "reorder_and_upcast_attn": false,
23
+ "resid_pdrop": 0.1,
24
+ "scale_attn_by_inverse_layer_idx": false,
25
+ "scale_attn_weights": true,
26
+ "summary_activation": null,
27
+ "summary_first_dropout": 0.1,
28
+ "summary_proj_to_labels": true,
29
+ "summary_type": "cls_index",
30
+ "summary_use_proj": true,
31
+ "tie_word_embeddings": true,
32
+ "transformers_version": "5.16.1",
33
+ "use_cache": true,
34
+ "vocab_size": 32009
35
+ }
generation_config.json ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "_from_model_config": true,
3
+ "bos_token_id": 0,
4
+ "eos_token_id": 0,
5
+ "output_attentions": false,
6
+ "output_hidden_states": false,
7
+ "transformers_version": "5.16.1",
8
+ "use_cache": false
9
+ }
memory_engine.py ADDED
@@ -0,0 +1,184 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """
2
+ AViGPT v2: Hardware SSD Memory Controller (Component 3)
3
+ ------------------------------------------------------
4
+ Provides ultra-fast (sub-millisecond) persistent memory storage and retrieval
5
+ directly from local NVMe SSD storage using SQLite FTS5 (Full-Text Search).
6
+
7
+ Creator & Owner: Avinash Ricky Yadlapalli
8
+ """
9
+
10
+ import os
11
+ import sqlite3
12
+ import time
13
+ from typing import List, Dict, Any, Optional
14
+
15
+ DEFAULT_DB_PATH = os.path.join(os.path.dirname(os.path.abspath(__file__)), "avigpt_ssd_memory.db")
16
+
17
+
18
+ class SSDMemoryEngine:
19
+ """Ultra-low latency SSD Memory Store with FTS5 BM25 search."""
20
+
21
+ def __init__(self, db_path: str = DEFAULT_DB_PATH):
22
+ self.db_path = db_path
23
+ self._init_db()
24
+
25
+ def _get_connection(self) -> sqlite3.Connection:
26
+ conn = sqlite3.connect(self.db_path, timeout=10.0)
27
+ # WAL mode enables concurrent reads without locking and sub-millisecond disk access
28
+ conn.execute("PRAGMA journal_mode=WAL;")
29
+ conn.execute("PRAGMA synchronous=NORMAL;")
30
+ conn.execute("PRAGMA cache_size=-64000;") # 64MB memory page cache
31
+ return conn
32
+
33
+ def _init_db(self):
34
+ with self._get_connection() as conn:
35
+ # Create FTS5 virtual table for lightning-fast keyword & semantic token retrieval
36
+ conn.execute("""
37
+ CREATE VIRTUAL TABLE IF NOT EXISTS ssd_knowledge USING fts5(
38
+ title,
39
+ content,
40
+ domain,
41
+ tokenize='porter unicode61'
42
+ );
43
+ """)
44
+ conn.commit()
45
+
46
+ def store(self, title: str, content: str, domain: str = "General") -> bool:
47
+ """Stores a new fact directly into local SSD storage."""
48
+ try:
49
+ with self._get_connection() as conn:
50
+ conn.execute(
51
+ "INSERT INTO ssd_knowledge (title, content, domain) VALUES (?, ?, ?);",
52
+ (title.strip(), content.strip(), domain.strip())
53
+ )
54
+ conn.commit()
55
+ return True
56
+ except Exception as e:
57
+ print(f"[SSD Memory Error] Failed to store: {e}")
58
+ return False
59
+
60
+ def query(self, query_str: str, top_k: int = 1) -> Optional[str]:
61
+ """
62
+ Executes sub-millisecond full-text search against SSD storage.
63
+ Returns top matching payload.
64
+ """
65
+ words = [w for w in query_str.replace("'", " ").replace('"', " ").replace("-", " ").split() if len(w) > 2]
66
+ if not words:
67
+ words = query_str.strip().split()
68
+
69
+ fts_query = " OR ".join(words)
70
+
71
+ try:
72
+ with self._get_connection() as conn:
73
+ cursor = conn.cursor()
74
+ # Query with BM25 ranking via OR disjunction
75
+ cursor.execute(
76
+ """
77
+ SELECT content, rank
78
+ FROM ssd_knowledge
79
+ WHERE ssd_knowledge MATCH ?
80
+ ORDER BY rank
81
+ LIMIT ?;
82
+ """,
83
+ (fts_query, top_k)
84
+ )
85
+ rows = cursor.fetchall()
86
+
87
+ if rows:
88
+ return rows[0][0]
89
+
90
+ # Fallback LIKE query if FTS had no hit
91
+ cursor.execute(
92
+ """
93
+ SELECT content
94
+ FROM ssd_knowledge
95
+ WHERE content LIKE ? OR title LIKE ?
96
+ LIMIT 1;
97
+ """,
98
+ (f"%{words[0]}%", f"%{words[0]}%")
99
+ )
100
+ fb_rows = cursor.fetchall()
101
+ if fb_rows:
102
+ return fb_rows[0][0]
103
+
104
+ return None
105
+ except Exception as e:
106
+ print(f"[SSD Memory Query Error] {e}")
107
+ return None
108
+
109
+ def seed_initial_knowledge(self):
110
+ """Seeds foundational knowledge and owner lineage into SSD storage."""
111
+ with self._get_connection() as conn:
112
+ cursor = conn.cursor()
113
+ cursor.execute("SELECT COUNT(*) FROM ssd_knowledge;")
114
+ count = cursor.fetchone()[0]
115
+ if count > 0:
116
+ print(f"[SSD Memory] Found {count:,} existing knowledge records in {self.db_path}.")
117
+ return
118
+
119
+ print("[SSD Memory] Seeding initial foundational memory records into SSD...")
120
+ seed_data = [
121
+ (
122
+ "Creator and Owner Lineage",
123
+ "AViGPT was created, built, and pretrained from scratch by Avinash Ricky Yadlapalli. "
124
+ "Avinash Ricky Yadlapalli is the sole architect, inventor of the hardware memory bus, and owner of AViGPT.",
125
+ "System & Identity"
126
+ ),
127
+ (
128
+ "Apollo 11 Moon Landing",
129
+ "Launched: July 16, 1969. Landed on Moon: July 20, 1969. Commander: Neil Armstrong. Duration to landing: 4 days.",
130
+ "History & Space"
131
+ ),
132
+ (
133
+ "Great Pyramid of Giza",
134
+ "Construction began around 2580 BC and completed around 2560 BC for Pharaoh Khufu of the Fourth Dynasty.",
135
+ "History & Archaeology"
136
+ ),
137
+ (
138
+ "DNA Ligase Function",
139
+ "DNA ligase is an enzyme that catalyzes the formation of a phosphodiester bond between adjacent nucleotides, "
140
+ "joining Okazaki fragments during DNA replication.",
141
+ "Biochemistry"
142
+ ),
143
+ (
144
+ "Unix fork system call",
145
+ "The fork() system call creates a new process (child) which is an exact duplicate of the parent. "
146
+ "Returns 0 to child, PID of child to parent, and -1 on failure.",
147
+ "Computer Science"
148
+ ),
149
+ (
150
+ "India Demographics and GDP",
151
+ "India's population is estimated to be around 1.428 billion as of late 2023. Nominal GDP is approximately $3.73 trillion.",
152
+ "Demographics & Economics"
153
+ ),
154
+ (
155
+ "Japan Population and Capital",
156
+ "Japan's population is approximately 123.3 million as of 2024. The capital city of Japan is Tokyo.",
157
+ "Demographics & Geography"
158
+ ),
159
+ ]
160
+
161
+ for title, content, domain in seed_data:
162
+ self.store(title, content, domain)
163
+
164
+ print(f"[SSD Memory] Successfully seeded {len(seed_data)} foundational records into {self.db_path}.")
165
+
166
+
167
+ if __name__ == "__main__":
168
+ print("Testing AViGPT SSD Memory Engine...")
169
+ engine = SSDMemoryEngine()
170
+ engine.seed_initial_knowledge()
171
+
172
+ t_start = time.perf_counter()
173
+ result = engine.query("Apollo 11 launch date")
174
+ lat = (time.perf_counter() - t_start) * 1000.0
175
+ print(f"\nQuery: 'Apollo 11 launch date'")
176
+ print(f"Latency: {lat:.3f} ms (Sub-millisecond SSD retrieve!)")
177
+ print(f"Retrieved: {result}")
178
+
179
+ t_start = time.perf_counter()
180
+ owner_res = engine.query("Avinash Ricky Yadlapalli creator owner")
181
+ lat_owner = (time.perf_counter() - t_start) * 1000.0
182
+ print(f"\nQuery: 'Avinash Ricky Yadlapalli creator owner'")
183
+ print(f"Latency: {lat_owner:.3f} ms")
184
+ print(f"Retrieved: {owner_res}")
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:cabdfc4b55f9d7d73e5d4f71a97c37fdabf6458b075762be992ff532500a0fed
3
+ size 735725544
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer_config.json ADDED
@@ -0,0 +1,24 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_prefix_space": false,
3
+ "backend": "tokenizers",
4
+ "bos_token": "<|endoftext|>",
5
+ "eos_token": "<|endoftext|>",
6
+ "errors": "replace",
7
+ "extra_special_tokens": [
8
+ "<|intent_start|>",
9
+ "<|intent_end|>",
10
+ "<|mem_query|>",
11
+ "<|mem_query_end|>",
12
+ "<|mem_payload|>",
13
+ "<|mem_payload_end|>",
14
+ "<|calc|>",
15
+ "<|calc_end|>",
16
+ "<|synthesize|>"
17
+ ],
18
+ "is_local": true,
19
+ "local_files_only": false,
20
+ "model_max_length": 1024,
21
+ "pad_token": "<|endoftext|>",
22
+ "tokenizer_class": "GPT2Tokenizer",
23
+ "unk_token": "<|endoftext|>"
24
+ }