DOMAIN SFT DATA GENERATION
-- / 7,000
--%
Waiting
Updated
RUN PROGRESS
-- / --
Updated
Loss
Train and validationThroughput
Tokens per secondOptimization
Learning rate and gradient normMODEL PROFILE
Miyashita-Lab LLM 300M Base
Built for full-context pretraining on a 10.64B-token Japanese corpus, with grouped-query attention and tied embeddings.
- Architecture
- --
- Depth
- --
- Hidden size
- Attention
- --
- Vocabulary
- --
- Precision
- --
TRAINING DATA
Educational Japanese web corpus
FineWeb2 Japanese text selected for educational quality, using the sample_10BT subset. Documents are whitespace-normalized, tokenized with the project SentencePiece model, and separated by EOS.
Dataset card- Dataset
- --
- Subset
- --
- Train documents
- --
- Train tokens
- --
- Validation
- --
- Revision
- --
Model--
Context--
Effective batch--
Dataset--
Hardware--