We are the Alibaba Infrastructure Operations Team, a critical component of a leading global technology enterprise's infrastructure strategy. Our mission is to ensure the highly efficient, stable, and secure operation of infrastructure across the globe. We are dedicated to delivering "last mile customer value" from our operational facilities, providing a seamless, reliable, and highly efficient service experience.
The Controls team serves as the regional technical capability center for infrastructure automation and monitoring systems. We bridge global headquarters engineering with regional deployment operations, owning the technical standards, platform development, and integration architecture for Building Management Systems (BMS), Electrical Power Monitoring Systems (EPMS), and Environmental Monitoring Systems (EMS). Our mission is to ensure that all infrastructure's automation and monitoring meets Alibaba's global reliability standards through rigorous technical governance, code-level quality assurance, and systematic knowledge transfer to regional deployment teams.
We operate as a matrix organization combining both infra operations engineers and platform development engineers, enabling us to address complex technical challenges that span infrastructure systems, software platforms, and regional deployment contexts. We actively seek engineers who can bring proven practices in observability, automation, and incident management to the infra controls domain — accelerating the evolution of our monitoring and automation capabilities through cross-domain knowledge migration.
- Own regional technical standards for BMS/EPMS/EMS integration, defining monitoring specifications, protocol requirements, and data quality benchmarks to ensure standardized, reliable monitoring coverage across all regional infrastructure.
- Lead EMS integration projects across regional colocation providers through structured program management. Coordinate cross-functional resources (internal engineering, colo provider teams, platform vendors) to ensure standardized telemetry access, data quality, and alarm management aligned with Alibaba global EMS standards.
- Develop integration tools, data validation scripts, and platform components to automate monitoring deployment, configuration, and quality assurance. Apply cloud-native practices (IaC, CI/CD, observability frameworks) to build and maintain the regional integration toolchain, improving delivery efficiency and reducing manual errors.
- Serve as regional technical escalation point for complex BMS/EPMS/EMS integration faults. Lead root cause analysis using structured methodologies (post-incident review, blameless retrospectives), develop permanent corrective actions, and establish knowledge base entries to prevent recurrence across the regional fleet.
- Review and validate electrical and HVAC automation control logic implemented by colocation providers. Orchestrate technical experts and vendor resources to ensure control strategies meet reliability, efficiency, and safety standards before production deployment.
- Conduct systematic technical risk assessment of regional automation infrastructure, including single points of failure analysis, redundancy validation, and mitigation roadmap development. Drive risk closure through structured follow-up with colo providers and internal stakeholders.
- Establish regional technical training programs for controls deployment engineers and local FM/FE teams. Transfer integration methodologies, debugging techniques, and operational best practices to build sustainable regional technical capability.
- Collaborate with global headquarters on platform roadmap, standards evolution, and tool development. Represent regional perspectives in global architecture reviews and ensure global policies are adapted to local infrastructure contexts.
Job Requirements
"Job Responsibilities:
- Own regional technical standards for BMS/EPMS/EMS integration, defining monitoring specifications, communication protocol requirements, and data quality benchmarks. Ensure all regional infrastructure achieve standardized, reliable monitoring coverage.
- Lead EMS integration projects across regional colocation providers through structured program management. Coordinate cross-functional resources (internal engineering teams, colo provider technical staff, platform vendors) to ensure standardized telemetry access, data quality, and alarm management aligned with Alibaba global EMS standards.
- Develop integration tools, data validation scripts, and platform components to automate monitoring system deployment, configuration, and ongoing quality assurance. Apply cloud-native engineering practices (Infrastructure as Code, CI/CD pipelines, observability frameworks) to build and maintain the regional integration toolchain, improving delivery efficiency and reducing manual errors.
- Serve as regional technical escalation point for complex BMS/EPMS/EMS integration faults. Lead root cause analysis using structured methodologies (post-incident review, blameless retrospectives), develop permanent corrective actions, and establish knowledge base entries to prevent recurrence across the regional fleet.
- Review and validate electrical and HVAC automation control logic implemented by colocation providers. Orchestrate technical experts and vendor resources to ensure control strategies meet reliability, efficiency, and safety standards before production deployment.
- Conduct systematic technical risk assessment of regional automation infrastructure, including single points of failure analysis, redundancy validation, and mitigation roadmap development. Drive closure of identified risks through structured follow-up with colo providers and internal stakeholders.
- Establish regional technical training programs for controls deployment engineers and local FM/FE teams. Transfer integration methodologies, debugging techniques, and operational best practices to build sustainable regional technical capability.
- Collaborate with global headquarters on platform roadmap, standards evolution, and tool development. Represent regional technical perspectives in global architecture reviews and ensure global policies are adapted to local infrastructure contexts. Drive continuous improvement by introducing cloud reliability engineering practices (SLO/SLI management, error budgets, chaos engineering principles) into infra operations where applicable."