Aarki logo

Aarki

Machine Learning Engineer, Infra (MLOps) - CN at Aarki

Beijing, ChinaFull-timeMachine LearningPosted 2 months ago
Apply with Pipeline

About the Role

<h2>Who Are We?</h2> <p>RZR is an AI-native advertising platform built for the next era of performance marketing. We operate at the intersection of machine learning, programmatic media, and full-funnel mobile growth, powering campaigns for some of the world's most ambitious advertisers. Our platform is purpose-built to deliver outcomes at scale, not just impressions.</p> <p>We are a team of builders, operators, and technologists who believe the advertising industry is overdue for a fundamental rethink. We move fast, operate with a high degree of ownership, and hold ourselves to an exceptionally high standard of craft.</p> <p>RZR is scaling aggressively with an active M&amp;A pipeline and a platform vision that puts us on a path to becoming an industry leader. This is a rare opportunity to join a company at an inflection point and help shape what it becomes.</p> <hr> <h2>Role Overview</h2> <p>As Machine Learning Engineer (Infra / MLOps) at RZR, you will design, build, and operate the model training and deployment infrastructure that powers our Demand-Side Platform (DSP). This role focuses on building scalable, flexible, and reliable systems for training models on billions of records across bidding, ranking, pacing, and fraud use cases.</p> <p>You will work at the intersection of machine learning, data platforms, and infrastructure — with a strong focus on automation, reproducibility, and reliability. This is a P0 priority hire directly tied to accelerating RZR's migration from legacy model training systems to Prefect-based DNN pipelines, enabling 100% UA on DNN.</p> <p>The right person for this role combines production-grade ML systems experience with a strong bias to automate, document, and build for reliability — someone who takes end-to-end ownership from data to serving, and is energized by the complexity of high-QPS real-time bidding infrastructure.</p> <h2>Key Responsibilities</h2> <ul> <li>Own the development and evolution of infrastructure that enables faster, more reliable, and more cost-efficient model training</li> <li>Design, build, and maintain automated model training and orchestration pipelines that scale across large datasets and support rapid recovery from failures</li> <li>Develop standardized training workflows that support experimentation, reproducibility, versioning, and traceability</li> <li>Build and operate observability and monitoring systems to detect data quality issues, training instabilities, model anomalies, and performance regressions</li> <li>Improve the efficiency, scalability, and maintainability of the model training codebase, defining and enforcing best practices across the ML organization</li> <li>Apply DevOps and MLOps best practices to machine learning training workflows, including CI/CD and automated testing</li> <li>Design, develop, and continuously optimize ML infrastructure for advertising recommendation systems, covering model training, online inference, model serving, and feature pipelines</li> <li>Build a high-performance, highly scalable ML platform to support rapid iteration and stable deployment of advertising recommendation models</li> <li>Optimize distributed training, online inference, and resource scheduling to continuously improve system performance, stability, and resource utilization</li> <li>Collaborate closely with algorithm engineers to drive efficient implementation of recommendation, ranking, and ad-serving models</li> <li>Stay current with advancements in ML infrastructure and AI technologies, including the application of LLMs in recommendation and advertising scenarios</li> </ul> <h2>Required Skills and Experience</h2> <h3>Must-Have</h3> <ul> <li>Strong proficiency in Python and Spark for ML training and deployment workflows</li> <li>Experience building and operating machine learning pipelines in production environments</li> <li>Hands-on experience with DevOps practices including CI/CD, infrastructure as code, and automated testing</li> <li>Experience with workflow orchestration tools such as Airflow or Prefect for ML pipelines</li> <li>Solid understanding of ML experimentation, reproducibility, model versioning, and dataset management</li> <li>Experience with large-scale data pipelines, feature generation, and offline/online data consistency</li> <li>Experience developing recommendation systems, advertising systems, search systems, or machine learning platforms</li> <li>Familiarity with mainstream ML frameworks such as PyTorch and TensorFlow</li> <li>Experience with ML infrastructure, model training, online inference, or model serving</li> </ul> <h3>Nice-to-Have</h3> <ul> <li>Familiarity with system programming languages including C++ and Rust</li> <li>Strong grasp of probability, statistics, and data analysis principles</li> <li>Exposure to online inference systems, gRPC/REST model endpoints, or streaming features via Kafka or Flink</li> <li>Ad-tech familiarity: auction dynamics, pacing, fraud signals, creative personalization</li> <li>Experience with large-scale distributed training, high-performance computing (HPC), or GPU optimization</li> <li>Familiarity with distributed computing frameworks such as Kubernetes, Ray, Spark, and Flink</li> <li>Interest in or practical experience with LLMs and their application in recommendation and advertising scenarios</li> <li>Experience with on-prem deployments of open source tools including Spark, ClickHouse, and Redash</li> <li>Strong English reading and writing skills for collaboration with global teams</li> </ul> <hr> <h2>Why Join RZR?</h2> <ul> <li><strong>End-to-end ownership across the full ML stack</strong> — data, features, training, evaluation, serving, A/B testing, and monitoring. You will not be working on one slice of the pipeline; you will shape all of it.</li> <li><strong>Real-time bidding and training pipelines at genuine scale</strong> — high QPS with tight latency SLOs. The infrastructure challenges here are not academic.</li> <li><strong>Shape the MLOps platform from the ground up</strong> — you will drive observability, data and model quality systems, and the MLflow-first platform, with direct influence on how the ML organization operates.</li> <li><strong>Mentorship and structured growth</strong> — paired with a senior ML engineer, with structured growth goals and a strong code review culture.</li> <li><strong>Immediate, measurable impact</strong> — your work will directly accelerate model iteration speed, improve feature quality, and improve offline/online metric alignment for RZR's core bidding and ranking systems.</li> <li><strong>Exposure to emerging AI technologies</strong> — RZR is actively exploring LLM applications in recommendation and advertising, and this role sits at the center of that work.</li> </ul> <hr> <h2>RZR Behaviors</h2> <p>RZR operates by eight core behaviors: Extreme Ownership · Move Fast · Drive for Excellence · Proactive Communication · Courage · Curiosity · Deliver Results · Manage Ambiguity</p>