Norconsult Telematics — الرياض - منطقة الرياض — دوام كامل
Job Posting Location : Riyadh, Saudi Arabia
Position Objective:
Bridge development and operations so that enterprise applications and IT infrastructure are highly available, secure, scalable, and delivered through reliable automation of IT infrastructure and enterprise applications. The role sits at the intersection of engineering velocity and operational excellence, improving how systems are built, deployed, monitored, and maintained. Intended for a practitioner who can translate platform and infrastructure needs into resilient, repeatable operational outcomes. Connect software delivery and operational reliability in a way that strengthens uptime, security, scalability, and automated delivery across enterprise applications and infrastructure. The scope spans infrastructure operations, deployment automation, reliability enablement, and the continuous improvement of environments and delivery processes that support development teams and business-critical systems. This role exists because the organization needs a senior operator-builder who can reduce manual operational overhead, improve consistency across environments, and create dependable delivery mechanisms that support growth without compromising control or stability. In team context, this person will work across development and operations boundaries, aligning engineering practices with operational needs and helping ensure that infrastructure and application delivery are both efficient and resilient.
Job Description & Responsibilities:
• Own outcomes that improve the availability of enterprise applications and supporting infrastructure by strengthening operational reliability, reducing preventable downtime, and ensuring that production environments remain stable under normal and peak conditions.
• Drive secure and automated delivery outcomes by improving the mechanisms through which infrastructure and enterprise applications are built, tested, released, and changed, reducing manual intervention and increasing repeatability.
• Improve scalability across platforms and environments by evolving infrastructure patterns and operational practices so that systems can support growth in usage, complexity, and business demand without disproportionate increases in operational effort.
• Strengthen the connection between development and operations teams by establishing workflows, handoffs, and technical guardrails that enable faster delivery while preserving system integrity, operational visibility, and accountability.
• Increase the security posture of infrastructure and delivery processes by embedding operational controls, secure configuration practices, and disciplined change management into day-to-day engineering and platform workflows.
• Reduce operational risk by identifying fragile processes, single points of failure, and environment inconsistencies, then driving remediation efforts that improve resilience and recovery readiness.
• Improve observability and operational decision-making by ensuring that systems produce the telemetry, alerts, and service-level signals needed to detect issues early, respond effectively, and prioritize reliability investments.
• Enable development teams to deliver more efficiently by providing dependable infrastructure foundations, standardized deployment pathways, and operational support models that reduce friction across the software lifecycle.
• Drive continuous improvement in infrastructure and application operations by using incidents, recurring issues, and delivery bottlenecks as inputs for durable process, tooling, and architecture improvements.
• Contribute senior-level operational leadership by shaping best practices, influencing engineering standards, and helping teams make pragmatic trade-offs between speed, stability, security, and scale.
• Lead the deployment, health monitoring, and lifecycle management of Container Orchestration platforms (OpenShift/Kubernetes), ensuring alignment with the organization’s strategic objectives.
• Manage a portfolio of mixed OS environments (Linux and Windows), handling server troubleshooting, kernel tuning, and automated provisioning via Infrastructure as Code (Terraform, Ansible).
• Oversee public/hybrid cloud infrastructure management (OpenShift, Azure, AWS), optimizing cloud resources and ensuring high availability across zones.
• To develop and execute project plans for CI/CD pipeline optimizations, reducing build times via caching, parallel execution, and multi-stage builds.
• Implement GitOps methodologies and execute zero-downtime releases utilizing Blue/Green or Canary deployment strategies.
• Troubleshoot and resolve pipeline failures, dependency conflicts, and restore broken webhook integrations across the development toolchain (Git, Jira, Azure).
• Act as the primary responder for application "slowness," isolating root causes across disk I/O limits, CPU throttling, or network latency.
• Establish robust observability using APM tools (Dynatrace, Datadog), building custom Grafana/Kibana dashboards, and tuning alerts to reduce fatigue.
• Troubleshoot complex network connectivity, DNS, DHCP, and VPN issues using packet analysis tools (tcpdump, wireshark) and optimize Load Balancer configurations (HAProxy, Nginx).
• Identify, assess, and mitigate risks associated with IT infrastructure by applying emergency hotfixes for zero-day vulnerabilities (CVEs) and automating OS/application patch cycles.
• To ensure that the project follows organizational policies, industry standards, and regulatory requirements by enforcing security hardening (CIS Benchmarks, SELinux) and managing secret vaults.
• Oversee quality assurance processes to ensure that infrastructure projects meet or exceed technical and business requirements, including successful backup verification and database disaster recovery (DR) drills.
• Enforce FinOps best practices to monitor and control cloud and infrastructure budgets, ensuring that projects are delivered within financial constraints.
• Perform database performance tuning, collaborating with developers on SQL query optimization, index tuning, and managing connection pool limits to ensure system stability.
• Allocate and optimize computing resources across namespaces, adjusting Kubernetes Horizontal Pod Autoscaler (HPA) thresholds based on real-time consumption.
• To develop and implement change management strategies to ensure smooth transitions during infrastructure upgrades, platform migrations, and toolchain updates.
• Engage with key stakeholders to ensure alignment on project goals, evaluating organizational impact of infrastructure changes to minimize disruption and maximize value.
• Build and lead trainings and communication efforts to ensure that all stakeholders (Developers, QA, Security) are prepared for infrastructure changes.
• To drive continuous improvements in project management and operations practices, automating manual administrative toil with Python/Bash/PowerShell scripts.
• Monitor the latest trends, incorporating innovative technologies (e.g., Site Reliability Engineering practices, SLOs/Error Budgets) into project plans where appropriate
• To promote a culture of innovation and excellence within the team, adopting new tools and methodologies to modernize legacy environments.
Qualifications & Experience:
• Bachelor’s degree in information technology, Computer Science, Engineering, or a related field.
• Advanced degree or relevant certifications (e.g., CKA/CKAD, AWS Certified DevOps Engineer, RHCE).
• 10+ years of experience in IT, with a focus on managing large-scale infrastructure projects, DevOps, or SRE operations.
• Deep understanding of IT infrastructure components, including networking, data centers, cloud services, and cybersecurity.
• Expertise in container orchestration (OpenShift/Kubernetes), CI/CD pipelines (Jenkins/GitLab), and Configuration Management (Ansible/Terraform).
• Expertise in regulatory compliance, data privacy, and information security in IT infrastructure pr