TRAINING MONITOR

Miyashita-Lab LLM 300M Base

Connecting

-- / 7,000

--% Waiting

Updated

Paper conversations--Target 4,000
Cross-document--Target 2,000
Lab overview--Target 1,000
Questions ready--Validated threads
Retry attempts--Rejected outputs
Rate / ETA--Calculating

-- / --

--%

Updated

Training loss--Latest step
Validation loss--Pending
Throughput--50-step median
ETA--Calculating
Learning rate--Cosine schedule
Peak VRAM--Allocated

Loss

Train and validation

Throughput

Tokens per second

Optimization

Learning rate and gradient norm

Miyashita-Lab LLM 300M Base

Built for full-context pretraining on a 10.64B-token Japanese corpus, with grouped-query attention and tied embeddings.

Architecture
--
Depth
--
Hidden size
--
Attention
--
Vocabulary
--
Precision
--

Educational Japanese web corpus

FineWeb2 Japanese text selected for educational quality, using the sample_10BT subset. Documents are whitespace-normalized, tokenized with the project SentencePiece model, and separated by EOS.

Dataset card
Dataset
--
Subset
--
Train documents
--
Train tokens
--
Validation
--
Revision
--
Model--
Context--
Effective batch--
Dataset--
Hardware--