via Indeed · 9 septembre 2026 ·il y a 11 jours

Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference

Apple
Paris
Cette annonce provient de Indeed
Voir l'annonce originale ↗

We are the Foundation Model Inference team within Cloud OS and AI Inference organization. We are on a mission to build the most highly performant, secure and private inference stack that powers Siri AI, Apple Intelligence and Apps that are powered with the largest foundation models. Our systems serve billions of queries daily across Siri AI, Apple Intelligence, Apple Search, Apple Music, Apple TV, App Store, iMessage, Photos, Camera, Spotlight \& Safari, at remarkably low latency with every ounce of compute extracted from the hardware beneath them. We optimise language, vision, and speech models with billions of parameters using state\-of\-the\-art techniques and ship them at Apple scale. This is a rare opportunity to directly shape how AI reaches billions of people worldwide.

Description

You will work at the intersection of research and production, partnering closely with the Foundation Model Research team and our external partners to bring cutting\-edge model architectures from prototype to planetary\-scale deployment. You will own hard problems in inference efficiency, hardware/software codesign, systems architecture, and tooling, and help set the technical direction for the engineers around you. This role sits within CloudOS and Private Cloud Compute (PCC) \- Apple's purpose\-built, privacy\-preserving cloud infrastructure for AI workloads. PCC represents a first\-of\-its\-kind approach to running foundation models in the cloud with verifiable privacy guarantees, and CloudOS is the systems foundation that makes it possible. You will be building and optimising inference systems on top of this infrastructure, working closely with platform and security teams to deliver both performance and trust at scale.","responsibilities":"Partner with the Foundation Model Research team and our external partners to optimise inference for the latest model architectures across language, vision, and speech.

Design and ship production\-grade inference systems serving millions of customers in real time.

Build profiling tools and simulators to identify and resolve performance bottlenecks across different hardware configurations and use cases.

Drive technical decisions on high\-throughput, low\-latency serving at supercomputing scale.

Mentor and grow engineers across the organisation.

Preferred Qualifications

Hands\-on experience with LLM inference stacks.

Working knowledge of GPU or TPU programming concepts.

Experience building and operating high\-throughput services at large distributed scale.

Experience building productions systems in Go or Python.

Strong knowledge of deep learning architectures including Transformers, encoder/decoder models, and multimodal variants.

Experience with inference optimization frameworks such as TensorRT\-LLM, vLLM, SGLang, TGI, or Nvidia Triton Server.

MS in Computer Science, Machine Learning, Artificial Intelligence, Data Science, or a related field

Minimum Qualifications

Experience leading complex ambiguous Machine learning projects end to end.

Proficiency in PyTorch or JAX

Experience working with Inference frameworks

Experienced in Python / Rust / Go lang or similar programming languages

Proficiency in deploying applications on cloud platforms (AWS, GCP or equivalent) using K8S and docker.

Le marché pour ce type de poste

Offres similaires
1 441
postes Ingénierie à Paris
Temps plein
83%
des offres Ingénierie en France
Télétravail possible
3%
des offres Ingénierie
Apple

64 postes ouverts · Barcelona, Berlin, Cork, ENG, Linz +9

📊 Ingénierie · France
29 574
offres actives
3.3%
Remote
Ø 1d
Ø en ligne
Compétences les plus demandées
ExcelERPISOPythonAWSCI/CDSQLAzureAgileLean

Questions fréquentes

Combien d'offres Ingénierie sont disponibles à Paris ?
Actuellement 1 441 postes en Ingénierie à Paris sur AlmostHired, dans 480 entreprises différentes. Nos données sont mises à jour quotidiennement.
Est-ce que les postes Ingénierie offrent du télétravail ?
3% des offres Ingénierie en France permettent le télétravail, partiel ou total. Pour filtrer spécifiquement les postes en remote, utilisez AlmostHired.
Comment savoir si je corresponds à cette offre ?
Déposez votre CV — notre IA compare votre profil aux exigences du poste et vous donne un score de compatibilité précis, avec les compétences qui correspondent et celles qui manquent.