Reliability Engineer
Bright Vision Technologies is thrilled to open its doors to a truly exceptional individual for our Reliability Engineer role. Imagine a position where your expertise directly shapes the robustness of vast systems, preventing hiccups before they even start. If you are someone who thrives on building bulletproof infrastructure, constantly pushing for better performance and smoother operations, then this is an exciting opportunity waiting for you. We are searching for a seasoned professional, a true SRE, ready to embed strong software engineering principles deep into the heart of our infrastructure. This is your chance to really make a difference, minimizing operational toil and maximizing the availability of critical distributed systems, all while enjoying the flexibility of a fully remote setup across the United States.
Overview
Bright Vision Technologies stands as a leading technology consulting and software development company. We are dedicated to delivering state of the art cloud, AI, data, and enterprise solutions throughout the United States. Joining our team means becoming part of an established and well respected organization that offers tremendous career growth potential and a dynamic work environment.
- Job Title: Reliability Engineer
- Location: 100% Remote across the United States
- Position Type: Full time, Direct W2
- Salary Range: $75,000 to $95,000 Annually
- Experience Required: Candidates should possess at least 6 years of relevant professional experience.
- Sponsorship: We encourage applications from U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates. Please note, we are unable to sponsor new H-1B visa petitions for this position.
Key Responsibilities
As a Reliability Engineer at Bright Vision Technologies, you will be a vital link between development and operations, applying your deep software engineering knowledge to infrastructure challenges. Your core responsibilities will include:
- Ensuring the availability, performance, and operational excellence of large scale distributed systems in production environments.
- Proactively identifying and resolving complex production issues, implementing robust solutions, and conducting thorough root cause analyses.
- Developing and deploying automation tools and frameworks to reduce manual toil and significantly enhance operational efficiency.
- Collaborating closely with our development teams to integrate reliability best practices throughout the entire software development lifecycle.
- Designing and implementing comprehensive monitoring, alerting, and logging solutions to maintain optimal system health and visibility.
- Participating actively in on call rotations to provide essential support for critical production systems.
- Driving continuous improvement initiatives aimed at boosting system reliability, overall performance, and operational efficiency.
- Applying strong software engineering principles to effectively solve complex infrastructure and operations problems.
Requirements
We are seeking a highly skilled individual who embodies deep systems knowledge and a passion for reliability. Ideal candidates will possess:
- At least 6 years of professional experience as a Reliability Engineer, Site Reliability Engineer (SRE), or a similar role focused on large scale distributed systems.
- Proven expertise in designing, building, and operating highly available and performant cloud based infrastructure.
- Strong programming skills in languages such as Python, Go, Java, or Ruby.
- Extensive hands on experience with leading cloud platforms, preferably AWS, Azure, or GCP.
- Deep understanding of Linux operating systems and fundamental networking principles.
- Familiarity with containerization technologies like Docker and orchestration tools such as Kubernetes.
- Experience with infrastructure as code tools, including Terraform or Ansible.
- Proficiency in monitoring, logging, and tracing tools, for example Prometheus, Grafana, ELK stack, or Datadog.
- A solid grasp of incident management, post mortem processes, and effective root cause analysis methodologies.
- Excellent problem solving abilities coupled with a proactive approach to ensuring system reliability and stability.
What You'll Gain
Joining Bright Vision Technologies means more than just a job; it is an opportunity to grow and thrive. Here is what you can look forward to:
- The chance to work with cutting edge technologies across cloud, AI, and data solutions.
- A stimulating environment where innovation, continuous learning, and professional development are highly valued.
- Becoming part of a supportive, collaborative, and rapidly growing team dedicated to excellence.
- Significant opportunities for professional development and clear pathways for career advancement within our organization.
- The flexibility and convenience of a fully remote work model, fostering a healthier work life balance.
- A competitive salary package and comprehensive benefits designed to support your well being.
How to Apply
Click the apply button below to view the full job details and submit your application directly through the employer's official page.

