AvinashRicky commited on
Commit
4a14f61
·
verified ·
1 Parent(s): 42a0779

docs: add benchmark charts, social banner, and updated repo links

Browse files
Files changed (1) hide show
  1. README.md +37 -12
README.md CHANGED
@@ -31,13 +31,15 @@ widget:
31
 
32
  # AViGPT: 183M Parameter Core with Native NVMe Hardware Bus
33
 
 
 
34
  **AViGPT** is a 183-million parameter autoregressive language model designed and pretrained from scratch by **Avinash Ricky Yadlapalli**.
35
 
36
  Rather than increasing parameter count to memorize factual records inside dense neural weights, AViGPT separates syntactic reasoning from factual storage. It couples a compact 183M reasoning core with a dedicated local NVMe SSD hardware memory bus. The model emits explicit control tokens to pause inference, execute sub-millisecond SQLite FTS5 full-text lookups on local storage, inject verified records into context, and complete generations with verified factual precision.
37
 
38
- * **Author:** Avinash Ricky Yadlapalli ([@Avinashricky211](https://huggingface.co/Avinashricky211))
39
- * **Interactive Demo:** [AViGPT Studio Space](https://huggingface.co/spaces/Avinashricky211/AViGPT-Studio)
40
- * **Code Repository:** [GitHub - Avinashricky211/AViGPT](https://github.com/Avinashricky211/AViGPT)
41
 
42
  ---
43
 
@@ -56,7 +58,7 @@ Rather than increasing parameter count to memorize factual records inside dense
56
  | **Alignment Dataset** | 25,850 multi-step hardware trajectories |
57
  | **Storage Engine** | Local SQLite 3 FTS5 (WAL mode, normal synchronous disk writes) |
58
  | **SSD Latency** | 1.18 milliseconds (NVMe average read latency) |
59
- | **RAM Footprint** | ~0.4 GB (runs comfortably on free-tier CPU or edge devices) |
60
 
61
  ---
62
 
@@ -93,6 +95,25 @@ User Instruction
93
 
94
  ---
95
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
96
  ## Hardware & Efficiency Comparison
97
 
98
  | System | Parameters | Minimum Hardware | Retrieval Latency | Knowledge Updates |
@@ -114,7 +135,7 @@ You can load and query the model directly via the `transformers` library:
114
  import torch
115
  from transformers import AutoModelForCausalLM, AutoTokenizer
116
 
117
- model_id = "Avinashricky211/AViGPT"
118
 
119
  tokenizer = AutoTokenizer.from_pretrained(model_id)
120
  model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.float32)
@@ -141,17 +162,20 @@ print(tokenizer.decode(outputs[0], skip_special_tokens=False))
141
  To run AViGPT with active sub-millisecond SSD queries and arithmetic execution:
142
 
143
  ```bash
144
- git clone https://huggingface.co/Avinashricky211/AViGPT
145
- cd AViGPT
146
- pip install torch transformers
 
147
  ```
148
 
 
 
149
  ```python
150
  import torch
151
  from transformers import GPT2LMHeadModel, GPT2TokenizerFast
152
  from autonomous_bus import AutonomousHardwareBus
153
 
154
- model_id = "Avinashricky211/AViGPT"
155
  tokenizer = GPT2TokenizerFast.from_pretrained(model_id)
156
  model = GPT2LMHeadModel.from_pretrained(model_id).to("cuda" if torch.cuda.is_available() else "cpu")
157
 
@@ -186,11 +210,12 @@ print(f"Total Response Latency: {metrics['total_latency_s']:.2f} s")
186
  ## Citation
187
 
188
  ```bibtex
189
- @article{yadlapalli2026avigpt,
190
  title={AViGPT: Decoupling Neural Reasoning from Parametric Memory via a Sub-Millisecond Native NVMe Hardware Bus},
191
- author={Yadlapalli, Avinash Ricky},
192
  year={2026},
193
  journal={AViGPT Technical Report},
194
- howpublished={\url{https://huggingface.co/Avinashricky211/AViGPT}}
 
195
  }
196
  ```
 
31
 
32
  # AViGPT: 183M Parameter Core with Native NVMe Hardware Bus
33
 
34
+ ![AViGPT Banner](assets/social_preview.png)
35
+
36
  **AViGPT** is a 183-million parameter autoregressive language model designed and pretrained from scratch by **Avinash Ricky Yadlapalli**.
37
 
38
  Rather than increasing parameter count to memorize factual records inside dense neural weights, AViGPT separates syntactic reasoning from factual storage. It couples a compact 183M reasoning core with a dedicated local NVMe SSD hardware memory bus. The model emits explicit control tokens to pause inference, execute sub-millisecond SQLite FTS5 full-text lookups on local storage, inject verified records into context, and complete generations with verified factual precision.
39
 
40
+ * **Author:** Avinash Ricky Yadlapalli ([@AvinashRicky](https://huggingface.co/AvinashRicky))
41
+ * **Code Repository:** [GitHub - Avinashricky211/AviGPT](https://github.com/Avinashricky211/AviGPT)
42
+ * **License:** MIT Open Source License
43
 
44
  ---
45
 
 
58
  | **Alignment Dataset** | 25,850 multi-step hardware trajectories |
59
  | **Storage Engine** | Local SQLite 3 FTS5 (WAL mode, normal synchronous disk writes) |
60
  | **SSD Latency** | 1.18 milliseconds (NVMe average read latency) |
61
+ | **RAM Footprint** | ~0.4 GB (runs comfortably on CPU or edge devices) |
62
 
63
  ---
64
 
 
95
 
96
  ---
97
 
98
+ ## Empirical Benchmarks & Performance Charts
99
+
100
+ ### 1. Training Convergence Curve
101
+ Cross-entropy loss declined from 3.3698 to 0.6935 over 1,800 steps on an NVIDIA T4 GPU:
102
+
103
+ ![Training Loss Convergence](assets/charts/loss_convergence.png)
104
+
105
+ ### 2. External Retrieval Latency (NVMe Bus vs. Network RAG)
106
+ Direct NVMe storage retrieval operates in 1.18 milliseconds, compared to 500 to 1,500 milliseconds for network-based RAG architectures:
107
+
108
+ ![Memory Latency Comparison](assets/charts/memory_latency_comparison.png)
109
+
110
+ ### 3. Hardware Footprint Comparison
111
+ AViGPT operates with a 0.4 GB memory footprint, running entirely on consumer CPUs without requiring dedicated GPU accelerators:
112
+
113
+ ![Hardware Footprint Comparison](assets/charts/vram_footprint_comparison.png)
114
+
115
+ ---
116
+
117
  ## Hardware & Efficiency Comparison
118
 
119
  | System | Parameters | Minimum Hardware | Retrieval Latency | Knowledge Updates |
 
135
  import torch
136
  from transformers import AutoModelForCausalLM, AutoTokenizer
137
 
138
+ model_id = "AvinashRicky/AViGPT"
139
 
140
  tokenizer = AutoTokenizer.from_pretrained(model_id)
141
  model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.float32)
 
162
  To run AViGPT with active sub-millisecond SSD queries and arithmetic execution:
163
 
164
  ```bash
165
+ git clone https://github.com/Avinashricky211/AviGPT.git
166
+ cd AviGPT
167
+ pip install -r requirements.txt
168
+ streamlit run app.py
169
  ```
170
 
171
+ Or programmatically in Python:
172
+
173
  ```python
174
  import torch
175
  from transformers import GPT2LMHeadModel, GPT2TokenizerFast
176
  from autonomous_bus import AutonomousHardwareBus
177
 
178
+ model_id = "AvinashRicky/AViGPT"
179
  tokenizer = GPT2TokenizerFast.from_pretrained(model_id)
180
  model = GPT2LMHeadModel.from_pretrained(model_id).to("cuda" if torch.cuda.is_available() else "cpu")
181
 
 
210
  ## Citation
211
 
212
  ```bibtex
213
+ @article{Avinash2026avigpt,
214
  title={AViGPT: Decoupling Neural Reasoning from Parametric Memory via a Sub-Millisecond Native NVMe Hardware Bus},
215
+ author={Avinash Ricky Yadlapalli},
216
  year={2026},
217
  journal={AViGPT Technical Report},
218
+ howpublished={\url{https://huggingface.co/AvinashRicky/AViGPT}},
219
+ url={https://github.com/Avinashricky211/AviGPT}
220
  }
221
  ```