Back to results page
Apply now
Company
AllianzPlace(s)
ParisIntern AI Evaluation Engineering (f/m/d), Paris
Internship / student job
Assurance
IT / Computer Science
Offer archived at 23/07/2026
Allianz
We are a strong global community connected by shared values, varied experiences, and a common goal — to grow with purpose. Our People Attributes reflect the values we live by and the mindset we bring to every challenge, every day.
We truly care for our employees, their individual needs, and their aspirations. We are a thriving community of thinkers, innovators, and individuals who share our purpose: We secure our future.
Tasks
- Conduct a structured benchmark of GenAI evaluation frameworks, both open-source (Ragas, DeepEval, TruLens, Phoenix/Arize, ARES, promptfoo) and commercial SaaS platforms (LangSmith, Braintrust, Humanloop, Galileo, Azure AI Evaluation); Compare metrics coverage, ease of integration, cost, licensing, scalability, and enterprise readiness
- Curate a gold-standard evaluation dataset (queries, expected outputs, source documents, edge cases) across GenAI Hub's core features
- Implement evaluation pipelines using the top 2–3 frameworks, measuring faithfulness, answer relevance, context precision, context recall, hallucination rate, and task-specific quality metrics
- Produce a quality baseline report identifying strengths and weaknesses per feature and per search index, with cross-framework comparison of metric consistency
- Experiment with programmatic prompt optimization tools (DSPy, TextGrad, MIPRO) to automatically improve retrieval and generation quality against the established baseline
- Integrate Responsible AI considerations into the evaluation framework — assessing outputs for bias, toxicity, fairness, and content safety — and recommend guardrails and automated checks for production deployments
- Deliver a tool selection recommendation, a reusable evaluation harness, optimized prompt candidates, and a comparative benchmark report (before/after) with cost/quality/safety trade-off analysis.
Profile
- Currently pursuing M1/M2 in Data Science, Machine Learning, Computer Science, or a related field
- Solid Python programming skills (scripting, data manipulation, API integration)
- Foundational understanding of NLP/ML concepts (embeddings, language models, retrieval systems)
- Familiarity with evaluation methodology and metrics design
- Ability to synthesize findings from multiple tools and produce clear, structured written reports
- Curiosity for applied research, tooling evaluation, and emerging AI practices
- Working proficiency in English (French is a plus)
- No prior enterprise experience required.
Apply
Offer archived at 23/07/2026
These jobs might also interest you:

Fr
De
En




