Jobiglo

No results.

LLM Evaluator (Model Response Analyst)

Odixcity Consulting

Remote
Remote Mid 🇬🇧 English

Job description

About the role

We are looking for a detail‑oriented LLM Evaluator (Model Response Analyst) to assess and improve the performance of large language models. The role is fully remote and involves evaluating AI‑generated text for factual accuracy, coherence, safety, bias and alignment with defined guidelines.

Key responsibilities

  • Evaluate and rank model‑generated responses using complex rubrics covering factuality, coherence, safety, instruction‑following and creativity.
  • Compare multiple model outputs for the same prompt, select the preferred one and justify the choice.
  • Provide concise feedback to modeling and training teams on recurring failure patterns.
  • Design adversarial prompts to expose biased, harmful or insecure model behavior and help patch safety gaps.
  • Collaborate with QA to refine evaluation guidelines for ambiguous or edge‑case scenarios.
  • Participate in cross‑checking sessions to calibrate scoring standards and ensure inter‑rater reliability.
  • Analyze underperforming outputs, hypothesize root causes and flag novel model behaviors for research.

Required profile

  • Minimum 2 years of professional experience in computational linguistics, data analysis, technical writing, NLP quality assurance or cognitive science.
  • Bachelor’s degree in Computer Science or a related field.
  • Strong ability to craft prompts that test model limits and to explain why a response is good or bad.
  • Experience with Reinforcement Learning from Human Feedback (RLHF) data collection.
  • Proven track record of maintaining consistency across evaluation teams and conducting calibration sessions.
  • Understanding of data distribution, A/B testing concepts and experimental design for model comparison.

Required skills

  • Prompt engineering
  • Reinforcement Learning from Human Feedback (RLHF) data collection
  • Dataset sourcing, cleaning and annotation for LLM fine‑tuning or evaluation
  • A/B testing applied to AI models
  • Design of experiments to compare model versions

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Odixcity Consulting.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Why are you reporting this job?

Thank you for your report. We will review this job.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

💬 Chat with us on Telegram Chat on WhatsApp

Published 1 month ago

Expires 2 weeks from now

20 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

Odixcity Consulting