llm-eval Local LLM evaluation framework running quantized models on a single 8GB GPU via llama.cpp's OpenAI-compatible llama-server. Setup