Falcon H1R 7B
About 4.6 GB at Q4; roughly 32K context on an 8 GB GPU
A compact hybrid reasoning model that genuinely fits an 8 GB card and performs well on reasoning and math tasks.
Caveat: Treat it as a reasoning assistant rather than a long-horizon coding agent; every layer still carries KV-cache cost.