utsabdahal34/NepaliGPT-base

NepaliGPT decoder-only model exported as standard Transformers GPT-2 weights. Install transformers, sentencepiece, torch, and huggingface_hub for the example below. The additional native model.pt requires the source package from https://github.com/utsab345/Nepali_GPT2.

Usage

from huggingface_hub import hf_hub_download
import sentencepiece as spm
import torch
from transformers import AutoModelForCausalLM

repo = "utsabdahal34/NepaliGPT-base"
tokenizer = spm.SentencePieceProcessor(model_file=hf_hub_download(repo, "tokenizer.model"))
model = AutoModelForCausalLM.from_pretrained(repo)
ids = torch.tensor([[tokenizer.bos_id()] + tokenizer.encode("नेपाल एक सुन्दर")])
output = model.generate(ids, max_new_tokens=40, do_sample=False)
print(tokenizer.decode(output[0].tolist()))

For an instruction-tuned checkpoint, construct the prompt with nepali_gpt2.sft.format_prompt(instruction, context).

Architecture

{
  "vocab_size": 16000,
  "context_length": 512,
  "emb_dim": 512,
  "n_heads": 8,
  "n_layers": 8,
  "drop_rate": 0.1,
  "qkv_bias": false
}

Data

Nepali Wikipedia and OSCAR Nepali are described in the source Colab notebook. The raw corpus and held-out split are not included in this export; verify upstream licenses before redistribution.

Training

Base checkpoint supplied by the project owner from the Colab notebook. Architecture is GPT-2 style with 16,000 SentencePiece tokens, 512 context, 512 hidden width, 8 layers and 8 heads. The exported checkpoint does not contain optimizer state or independently verifiable training-step metadata.

Evaluation

The supplied base checkpoint scores 3/5 on the repository cloze smoke set. Full generation and CPU benchmark reports are stored in docs/measurements when available. This score is not a broad quality benchmark.

Limitations

This is a base language model, not instruction tuned. It may repeat text, hallucinate facts and reflect source-data bias. Held-out perplexity, multilingual baseline, human instruction-following scores and CUDA measurements are unavailable in this export.

Downloads last month
292
Safetensors
Model size
33.7M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support