Website Scopely

Apply for the Senior Site Reliability Engineer – Security position on the Gen AI team at Scopely (Bangalore). Build automated remediation loops, telemetry pipelines, and runtime reliability for production AI platforms.

About the Company: Scopely

Scopely is a global interactive entertainment giant and a premier force in mobile gaming, ranking as the #1 mobile games company in the United States and #2 globally. With over $10 billion in lifetime revenue, Scopely develops, publishes, and live-operates record-breaking cultural phenomena across mobile, web, PC, and console platforms—including massive blockbusters like MONOPOLY GO!, Pokémon GO, Stumble Guys, MARVEL Strike Force, and Star Trek™ Fleet Command.

This elite engineering position is embedded directly within the specialised Gen AI Team based in Bangalore, Karnataka, India. As a crucial part of Scopely’s centralised technology frontier, the Gen AI team architects, scales, and secures the foundational infrastructure powering internal agentic networks, large language model (LLM) workflows, and automated production runtime systems across global game operations.

About the Role: Senior SRE – Security (AI Infrastructure & Reliability)

Are you a code-first infrastructure engineer who believes security is an engineering problem solved with automated pipelines, telemetry, and self-healing systems rather than manual ticket triage? Scopely is looking for a Senior Site Reliability Engineer – Security to secure and harden its cutting-edge AI platform in Bangalore. This is not a generic SOC or compliance audit role. It is a highly technical, hands-on infrastructure engineering position designed for SREs or platform developers bringing 5 or more years of experience who excel at keeping production environments observable, controllable, and extraordinarily resilient.

In this role, you will take ownership of the runtime stability, observability, and automated hardening loops of Scopely’s internal generative AI and agentic systems. Your daily focus will involve writing clean Python automation scripts, deploying immutable cloud primitives via Terraform or Pulumi, and engineering distributed tracing mechanics (using correlation IDs and tool-call logs) to map agent interactions across service boundaries. By codifying detection policies, orchestrating proactive automated remediation loops, and establishing kill-switches and rollback drills, you will play a defining role in de-risking advanced Model Context Protocol (MCP) integrations and scaling secure AI infrastructure safely.

Key Responsibilities & AI Platform Workflows

  • AI Observability Layer Architecture: Design, build, and maintain comprehensive telemetry layers for LLM platforms, encompassing granular audit trails, tool-call event logs, unique correlation IDs, and deep distributed tracing across service boundaries.

  • Automated Remediation Engineering: Code and deploy automated “findings-to-fix” resolution loops, seamlessly ingestion telemetry alerts from Cloud Security Posture Management (CSPM) systems and converting them into programmable mitigation workflows.

  • System Hardening & Resilience Controls: Implement robust reliability guardrails for production AI agents, including specialized alerting threshold patterns, programmatic kill-switch configurations, strict rate limiting, and configuration drift detection.

  • Policy As Code & Toil Reduction: Minimize operational friction and eliminate regression vectors by codifying architectural guardrails, vulnerability scanning hooks, and compliance runtime checks directly into continuous delivery (CI/CD) pipelines.

  • Production Readiness & Architecture Reviews: Critically evaluate incoming infrastructure and AI application blueprints, focusing heavily on secure secrets management, risky external API calls, secure Model Context Protocol (MCP) telemetry, and operational failure surfaces.

  • Incident Response & Rollback Drills: Lead simulation exercises, rollback drills, and evidence-driven postmortems for runtime AI issues, building highly reproducible playbooks to streamline recovery workflows.

  • Cross-Functional Security Alignment: Function as the primary operational link for the AI platform, partnering closely with central IT, platform engineering, and SOC units during high-severity escalation windows.

Candidate Prerequisites & Technical Skill Stack

Successful candidates must present a production-tested background in site reliability engineering or platform operations, backed by a strong software engineering mentality and robust automation instincts.

Required Experience & Educational Baseline:

  • Professional Tenure: 5+ years of dedicated experience working within Site Reliability Engineering (SRE), Production Engineering, Platform Operations, DevSecOps, or dedicated Security Automation tracks.

  • Programming Dexterity: Exceptional software development fluency, specifically utilizing Python to interact with web APIs, restructure log schemas, and drive cloud automation scripts.

  • Architectural Philosophy: Proven expertise in incident handling, disaster recovery mechanisms, structural rollback strategies, and the design of high-signal Service Level Indicator (SLI) and Service Level Objective (SLO) baselines.

Core Technical Stack Proficiencies:

  • Infrastructure-as-Code (IaC): Advanced, hands-on experience building and managing cloud topography via Terraform or Pulumi.

  • Observability & Telemetry Systems: Profound structural understanding of modern distributed observability frameworks—mastering the ingestion, correlation, and parsing of complex metrics, structural logs, and distributed traces.

  • Cloud Platform Security Architecture: In-depth knowledge of Amazon Web Services (AWS) core ecosystems, including identity and access management (IAM) diagnostic scripting, VPC telemetry, cloud trail configurations, and data stream pipelines.

  • CI/CD Pipeline Integration: Mastery of modern pipeline execution platforms to embed dynamic application security testing (DAST), automated policy validations, and regression check loops.

Preferred Domain Pluses & Bonus Points:

  • Software Engineering Foundation: Prior experience as a backend software engineer, showcasing mature patterns for code reviews, unit testing, and systemic debugging.

  • Agentic AI Telemetry: Hands-on experience scaling agent runtimes, tracking prompt-token telemetry, or architecting internal platform guardrails for LLM-integrated products.

  • Enterprise Tool Integration: Familiarity with modern infrastructure security platforms (e.g., Wiz, CrowdStrike, Orca Security, AWS GuardDuty) or runtime monitoring tools like WAF/RASP (Runtime Application Self-Protection).

  • Compliance & Privacy: Familiarity with compliance-oriented logging frameworks or privacy-aware telemetry ingestion methods.

Core Position Specifications

  • Position Title: Senior Site Reliability Engineer – Security (Gen AI)

  • Hiring Organisation: Scopely

  • Corporate Work Location: Bangalore, Karnataka, India (On-Site Operational Base)

  • Experience Allotment: 5+ Years of core SRE / Platform Security experience

  • Employment Framework: Full-Time, Permanent Technical Assignment

  • Functional Focus: AI Agent Telemetry, Automated Remediation Loops, Policy-as-Code, Infrastructure Hardening, and Distributed Tracing

  • Industry Placement: Video Games / Interactive Entertainment / Generative AI Cloud Engineering

  • Functional Department: Gen AI Infrastructure & Platform Engineering

Upload your CV/resume or any other relevant file. Max. file size: 2 GB.