Empresa: Techrx
About the product We run an agentic marketplace for car buyers. Instead of forcing people through filter dropdowns, our users talk to an agent: “I need a 7 seater under R$120k that’s fuel-efficient and suitable for a family with two kids.” The agent researches, compares real inventory, explains trade-offs, and takes the person to a car they can actually buy. Behind that conversation sits a fleet of models — routers, retrievers, rankers, extractors, summarizers, judges - each doing a specific job under a latency and cost budget. This role owns that fleet. Responsibilities: Evaluation — the offline and online eval suites that tell us whether the agent is actually getting better: golden datasets from real traffic, calibrated LLM-as-judge rubrics, and eval as a release gate for every model, prompt, or pipeline change. Model strategy — deciding which model serves which step, and building the router that balances quality, latency, and cost. Benchmarking new releases against our own evals rather than vendor claims. Fine-tuning and adaptation — SFT, LoRA, and preference tuning on automotive-domain tasks when it genuinely beats better prompting or retrieval, plus the data flywheel that feeds it. The deep research pipeline — multi-step retrieval and synthesis across inventory, specs, pricing, reviews, and ownership cost, with grounding and citations we can trust. Serving and lightweight MLOps — inference services, embedding and index refresh jobs, versioning, staged rollout and rollback, and the observability to see quality and cost in production. Fallback and redundancy — multi-provider fallback chains, circuit breakers, and graceful degradation, so a slow or unavailable model never becomes a broken experience for the buyer. Cost-conscious infrastructure — owning cost per conversation and keeping the stack lean. Requisitos: Required 3+ years building and shipping ML or LLM systems in production (not just notebooks or POCs). Fine-tuning working experience ( LoRA/QLoRA, SFT, DPO ) and familiarity with serving stacks like vLLM, TGI, or Ollama . Working experience with agent frameworks and tool-calling architectures , and their failure modes. Experience with LLM observability tooling (Datadog LLM Observability, etc.). Multi-provider LLM operations and resilience patterns at real traffic volume. Strong Python ; comfortable owning a service end to end, not handing it off. Real, hands-on experience evaluating LLM systems — you’ve built an eval set, argued about a rubric, and caught a regression before users did. Practical experience with retrieval systems: embeddings, vector stores, hybrid search, reranking . Solid engineering fundamentals: APIs, async, testing, CI/CD, containers, cloud . Cost-awareness as an instinct. You reach for the cheapest thing that meets the bar. Advanced English for conversation with global teams . Nice to have Familiarity with tools and practices such as Unsloth, Hugging Face, MLflow, LangSmith, and/or Weights & Biases. Marketplace, e-commerce, recommender, or search ranking background. WORK MODEL: Hybrid. HIRING MODEL: CLT.