microsoft/phi-1
Update README.md (#15)
Upload data_summary_card.md (#14)
fix(config): Removes auto_map since it is not used anymore.
Update README.md
Delete modeling_phi.py
Delete configuration_phi.py
Update README.md
Delete pytorch_model.bin
Adding `safetensors` variant of this model (#9)
Update LICENSE
Update README.md
Update README.md
Update README.md
Update README.md
Update config.json
Update modeling_phi.py
Update README.md
Update README.md
Update modeling_phi.py
Update modeling_phi.py
Update modeling_phi.py
Upload modeling_phi.py
Delete Research License.docx
Upload 5 files
Update config.json
Update modeling_phi.py
Update modeling_phi.py
Update configuration_phi.py
fix(root): Fixes relative paths.
chore(root): Updates files to internal transformers implementation.
Update README.md
Upload 4 files
Update README.md
Update README.md
chore(readme): Updates with clear information.
Disables inference API to prevent mismatch with HF implementation.
fix(modeling_phi): Fixes initial generation with length larger than context length.
fix(modeling_phi): Fixes cached generation when above maximum context length.
Fixes exceeding maximum sequence length when using generate().
Uses native torch decorator for disabling autocast.
Adds disable_autocast support for different device types.
Fixes any potential overflow when calculating attention weights.
Delete modeling_mixformer_sequential.py
Delete configuration_mixformer_sequential.py
Upload pytorch_model.bin
Update to new model interface.
Improves type hinting on configuration arguments.
Fixes flash-attn import with a try/except statement
Adds support for flash-attn rotary embedding and fused dense layers.
Adds support for MQA/GQA and attention mask during training / fine-tuning.
