emrekuruu/job-search-distill
Viewer β’ Updated β’ 24.7k β’ 19
GGUF-quantized base model + two LoRA adapters for resume-aware job search. The
llama.cpp-runtime companion to emrekuruu/job-searcher-qwen3-8B
(the safetensors original). Distilled from DeepSeek V4 Pro onto Qwen3-8B.
| File | Size | Purpose |
|---|---|---|
Qwen3-8B-Q4_K_M.gguf |
~5 GB | Base model, q4_K_M quantized |
query_gen.lora.gguf |
~100 MB | LoRA: resume β LinkedIn search queries |
fit_eval.lora.gguf |
~100 MB | LoRA: (resume, job) β 5Γ20-pt fit score + reasoning |
llama-cpp-python)
from llama_cpp import Llama
from huggingface_hub import hf_hub_download
repo = "emrekuruu/job-search-gguf"
base = hf_hub_download(repo, "Qwen3-8B-Q4_K_M.gguf")
q_lora = hf_hub_download(repo, "query_gen.lora.gguf")
e_lora = hf_hub_download(repo, "fit_eval.lora.gguf")
llm = Llama(
model_path=base,
n_gpu_layers=-1,
n_ctx=16384,
lora_paths=[q_lora, e_lora],
lora_scales=[1.0, 0.0], # query_gen ON, fit_eval OFF
verbose=False,
)
# Switch tasks by re-scaling the adapters:
# llm.set_lora_scale(0, 0.0); llm.set_lora_scale(1, 1.0) # β fit_eval
resp = llm.create_chat_completion(
messages=[
{"role": "system", "content": "..."},
{"role": "user", "content": "..."},
],
response_format={"type": "json_object"},
max_tokens=4096,
)
print(resp["choices"][0]["message"]["content"])
Qwen/Qwen3-8B, quantized to q4_K_M via llama.cpp's convert_hf_to_gguf.py.emrekuruu/job-search-distill,
then converted with llama.cpp's convert_lora_to_gguf.py.emrekuruu/job-search-lora.Apache-2.0 (matches the Qwen3-8B base). Teacher labels used during training were generated via the DeepSeek API and are subject to DeepSeek's Open Platform Terms of Service.
4-bit