
6.1 Model-development lifecycle
A responsible model lifecycle includes several stages: Data selection and preparation. Define data sources, permitted uses, quality criteria, privacy constraints, and filtering procedures. Keep a record of dataset versions and processing steps. Tokenization and representation. Select and test the method used to convert input text or other modalities into model-ready representations. Tokenizer changes can affect compatibility and evaluation. Architecture and training. Choose a model architecture and training objective that match the available hardware, data, and intended use. Record the configuration, random seeds where applicable, software versions, and resource use. Checkpointing and recovery. Save model weights and training state in a controlled format. Test that checkpoints can be restored rather than assuming that a successful save means a usable checkpoint. Evaluation. Test capabilities, failure modes, robustness, safety, and performance on held-out data. Avoid relying on a single benchmark or cherry-picked demonstration. Inference and serving. Measure latency, throughput, memory use, stability, and cost under realistic workloads. Provide clear error handling and usage limits. Release and monitoring. Document the model’s intended use, known limitations, version, and evaluation results. Monitor failures and maintain a process for updates or withdrawal.6.2 Model variants
If Pulsar develops into a family of models, each variant should have a defined role. Possible categories include:- Compact models: lower resource requirements for local use or narrow tasks.
- General assistant models: broad language and reasoning support.
- Coding-focused models: assistance with programming tasks and repository context.
- Multimodal models: models that process more than text, if and when the relevant capabilities are implemented.
- Research checkpoints: experimental versions that are not intended for ordinary production use.
6.3 Evaluation philosophy
A model should be evaluated against the job it is expected to perform. Relevant dimensions may include:
Evaluation reports should describe the test set, scoring method, sample size where relevant, known limitations, and comparison conditions. Numbers without methodology are decoration.