Tung Tung SahurJobJobSahur
Wróć do wszystkich ofert
La Vaca Saturno Saturnita
Npv

Multimodal ML Engineer

Npv

Na pół etatusonstiges

Lokalizacja

Paris

Wynagrodzenie

Do negocjacji

Opublikowano

33 min temu

Opis

We're looking for a Multimodal ML Engineer to join White Circle, an AI Safety company building the policy enforcement and optimization layer for AI systems. Backed by $11M from senior leaders at OpenAI, Anthropic, HuggingFace, Mistral, and DeepMind, White Circle processes 100M+ API calls monthly and runs its own LLMs in production.

You will

  • Train and fine-tune large-scale multimodal models (vision-language, audio, speech, video) from scratch and from pretrained checkpoints.

  • Design experiments, build multimodal data pipelines, and train MoE architectures.

  • Build alignment pipelines (SFT, DPO, GRPO), optimize for production (quantization, distillation, streaming), and deploy end-to-end.

  • Define evaluation metrics that actually matter for the product.

Requirements

  • 3+ years training large-scale multimodal models.

  • Strong PyTorch and distributed training experience (DeepSpeed, FSDP).

  • Deep familiarity with multimodal architectures – LLaVA, Qwen-VL, InternVL, Audio Flamingo, Whisper, HuBERT, Conformer or similar.

  • Hands-on RLHF/alignment across modalities (GRPO, DPO, reward modeling).

  • Both audio and video experience required – sequence modeling for each, plus large-scale dataset curation and production inference optimization.

  • Relocation to Paris or London (hybrid) required.

Bonus

  • Audio signal processing fundamentals – spectrograms, mel features, noise reduction.

  • MoE architecture experience.

We offer

  • $100k–$250k/year salary + equity; higher figures can be negotiated.

  • Official employment, visa and relocation help.

Compensation: $100K – $250K • Higher figures and equity are negotiable

  • • $100K – $250K • Higher figures and equity are negotiable

Find more English Speaking Jobs in France on Arbeitnow

Zaimportowane z ArbeitNow · Zobacz oryginalne ogłoszenie

Wymagania

  • 3+ years training large-scale multimodal models.
  • Strong PyTorch and distributed training experience (DeepSpeed, FSDP).
  • Deep familiarity with multimodal architectures – LLaVA, Qwen-VL, InternVL, Audio Flamingo, Whisper, HuBERT, Conformer or similar.
  • Hands-on RLHF/alignment across modalities (GRPO, DPO, reward modeling).
  • Both audio and video experience required – sequence modeling for each, plus large-scale dataset curation and production inference optimization.
  • Relocation to Paris or London (hybrid) required.