Commit 5a41992

authored

convert-hf : support bfloat16 conversion (#7158)

* convert-hf : support bfloat16 conversion * gguf-py : flake8 fixes * convert-hf : add missing space after comma * convert-hf : get bit-exact same output as ./quantize The quantization version was missing. * convert-hf : don't round bf16 NANs * convert-hf : save some memory with np.int16 intermediate bf16 weights * convert-hf : more closely match llama.cpp with which weights to keep in f32 * convert-hf : add --outtype auto-f16 A reason for this to exist is for model quantizers who want an initial GGUF with the most fidelity to the original model while still using a 16-bit float type instead of 32-bit floats. * convert-hf : remove a semicolon because flake8 doesn't like it It's a reflex from when programming in C/C++, I guess. * convert-hf : support outtype templating in outfile name * convert-hf : rename --outtype auto-f16 to --outtype auto

1 parent fae9d23 commit 5a41992Copy full SHA for 5a41992

5 files changed

+406

-184

lines changed

convert-hf-to-gguf.py
gguf-py/gguf

5 files changed

+406

-184

lines changed

Comments

(0)

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Commit 5a41992

5 files changed

5 files changed

File tree

5 files changed

5 files changed

0 commit comments