GGUF
A file format (llama.cpp/Ollama) for running quantized models efficiently on CPU/Mac.
A file format (llama.cpp/Ollama) for running quantized models efficiently on CPU/Mac. (M08, M14)
A model file format (from llama.cpp) with flexible quant levels (Q4_K_M, etc.); popular for CPU/consumer/local use.