Nokia - Oracle TEST Logo

Nokia - Oracle TEST

Principal Architect/Lead- Observability

Posted 17 Hours Ago
Be an Early Applicant
In-Office or Remote
Hiring Remotely in United States
Senior level
In-Office or Remote
Hiring Remotely in United States
Senior level
Lead ownership of observability practices and instrumentation standards across distributed AI infrastructure. Manage the observability stack, ensure proper service instrumentation, build data pipelines (OLAP, Kafka, ETL, streaming), collaborate with SRE, platform, and UI/UX to create dashboards, explore graph/semantic enhancements, mentor teams, and drive adoption of industry tools like Prometheus, OpenTelemetry, Grafana, Tempo, and Victoria Metrics.
The summary above was generated by AI
We are seeking an experienced and visionary Principal Architect/Lead to take ownership of our observability practices and instrumentation standards. This role is crucial in ensuring we have a robust and efficient observability stack, enabling our engineering teams to monitor and optimize our distributed AI infrastructure effectively. The successful candidate will work closely with various teams, including Platform, SRE, UI, and UX, to establish best practices and drive the adoption of industry-leading observability tools and methodologies.

 
Responsibilities
  • Own and manage the observability stack, including tooling and instrumentation standards, across the entire codebase.
  • Work closely with SRE, engineering, and platform teams to embed observability practices from day one of a project's lifecycle.
  • Ensure proper instrumentation of services and push for high-quality observability data collection.
  • Stay up-to-date with the latest observability tooling and technologies, such as Prometheus, OpenTelemetry, Grafana, Tempo, and Victoria Metrics.
  • Collaborate with UI/UX designers to create intuitive and informative observability dashboards and visualizations.
  • Develop and maintain OLAP databases, Kafka, ETL pipelines, and streaming platforms to support observability data processing and analysis.
  • Explore and implement graph-databases, ontologies, and semantic layers to enhance observability and knowledge-base capabilities.
  • Provide expertise and guidance to engineering teams on best practices for distributed AI infrastructure observability.
  • Actively participate in code reviews and provide feedback to ensure proper instrumentation and observability practices are followed.
Qualifications
  • Expert knowledge of observability tooling, including Prometheus, OpenTelemetry, Grafana, Tempo, and Victoria Metrics.
  • Experience with OLAP databases, Kafka, ETL pipelines, and streaming platforms (Spark, Flink) is essential.
  • Must be hands-on in designing and implementing code; Experience with Kubernetes; Experience with cloud technologies
  • Familiarity with graph-databases, ontologies, semantic layers, and knowledgebases is highly advantageous.
  • Understanding of distributed systems and AI infrastructure at scale.
  • Ability to work closely with engineering teams and drive adoption of observability practices.
  • Excellent communication and collaboration skills to work effectively with cross-functional teams.
  • Experience in leading and mentoring a team of observability experts is desirable.
  • A proven track record of implementing and optimizing observability stacks in large-scale projects.
  • Strong problem-solving and analytical skills, with the ability to identify and resolve complex issues.
  • A passion for staying updated with the latest advancements in observability and monitoring technologies.
  • Master's degree in Computer Science, Engineering, or a related field; PhD preferred.
     
About Us
Advancing connectivity to secure a brighter world.

Nokia is a global leader in connectivity for the AI era. With expertise across fixed, mobile and transport networks, powered by the innovation of Nokia Bell Labs, we’re advancing connectivity to secure a brighter world. 

Learn more about life at Nokia.


Our recruitment process

We act inclusively and respect the uniqueness of people. Our employment decisions are made regardless of race, color, national or ethnic origin, religion, gender, sexual orientation, gender identity or expression, age, marital status, disability, protected veteran status or other characteristics protected by law. We are committed to a culture of inclusion built upon our core value of respect.

If you’re interested in this role but don’t meet every listed requirement, we still encourage you to apply. Unique backgrounds, perspectives, and experiences enrich our teams, and you may be just the right candidate for this or another opportunity.

The length of the recruitment process may vary depending on the specific role's requirements. We strive to ensure a smooth and inclusive experience for all candidates. Discover more about the recruitment process at Nokia. 

About the Team
Some of our benefits in US:
  • Corporate Retirement Savings Plan
  • Health and dental benefits
  • Short-term disability, and long-term disability
  • Life insurance, and AD&D – Company paid 2x base pay
  • Optional or Supplemental life and AD&D insurance (Employee/Spouse/Child)
  • Paid time off for holidays and Vacation
  • Employee Stock Purchase Plan
  • Tuition Assistance Plan
  • Adoption assistance
  • Employee Assistance Program/Work Life Resource Program

The above benefits exclude students.


Disclaimer for US/Canada

Nokia Maintains broad annual base salary ranges for its roles in order to account for variations in knowledge, skills, experience and market conditions, and with consideration to internal peer equity. Check the salary ranges in the job info section for this role.

All North America job posts will post for a minimum of 7 calendar days and up to 180 days or until candidate/s identified.

Similar Jobs at Nokia - Oracle TEST

17 Hours Ago
In-Office or Remote
United States
Senior level
Senior level
Information Technology
Lead design and operation of a massively distributed edge compute platform using hyperscaler services and Kubernetes. Build external APIs, optimize high-throughput L7 traffic routing, ensure fault tolerance and security, present scalability metrics, mentor engineers, run design reviews, and collaborate across AI/ML, hardware, network, and security teams.
Top Skills: APIsAWSAzureContainerizationDistributed SystemsDockerGCPKubernetesL7 Traffic RoutingNetwork Protocols
17 Hours Ago
In-Office or Remote
United States
Senior level
Senior level
Information Technology
Lead and grow a QA team, design and implement test strategy with E2E automation and scale testing, build test automation, run scale tests with device simulators and traffic generators, apply chaos engineering to validate distributed systems resilience, integrate testing with development, analyze results, and mentor team members.
Top Skills: Chaos EngineeringCloudDevice SimulatorsEnd-To-End (E2E) Test AutomationKubernetesNokia Quality ProcessesRegression TestingScale TestingTest Automation FrameworksTraffic Generators
17 Hours Ago
In-Office or Remote
United States
Senior level
Senior level
Information Technology
Develop and maintain test automation frameworks, execute end-to-end automated tests, integrate automation into the SDLC, perform resilience/chaos testing, manage test environments (cloud and remote hardware), analyze results, and mentor junior engineers to improve test quality and efficiency.
Top Skills: AIChaos EngineeringCloud PlatformsDistributed SystemsKubernetesRemote Hardware ManagementTest Automation Frameworks

What you need to know about the Charlotte Tech Scene

Ranked among the hottest tech cities in 2024 by CompTIA, Charlotte is quickly cementing its place as a major U.S. tech hub. Home to more than 90,000 tech workers, the city’s ecosystem is primed for continued growth, fueled by billions in annual funding from heavyweights like Microsoft and RevTech Labs, which has created thousands of fintech jobs and made the city a go-to for tech pros looking for their next big opportunity.

Key Facts About Charlotte Tech

  • Number of Tech Workers: 90,859; 6.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lowe’s, Bank of America, TIAA, Microsoft, Honeywell
  • Key Industries: Fintech, artificial intelligence, cybersecurity, cloud computing, e-commerce
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (CED)
  • Notable Investors: Microsoft, Google, Falfurrias Management Partners, RevTech Labs Foundation
  • Research Centers and Universities: University of North Carolina at Charlotte, Northeastern University, North Carolina Research Campus

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account