Already filled

Don't miss the next one. Get matching roles delivered to your inbox.

SU

Sundayy

ML Platform Engineer

Job summary

United States
Software Developer

Work model

Fully remote
Only US
3 weeks ago
Job description

About The Company

Bright Vision Technologies is a forward-thinking software development company dedicated to building innovative solutions that help businesses automate and optimize their operations. We leverage cutting-edge technologies to create scalable, secure, and user-friendly applications. Our focus is on delivering high-quality, reliable, and efficient software solutions that empower organizations to achieve their strategic goals. As a growing organization, we value talent, innovation, and a collaborative approach to solving complex technological challenges. Our team comprises experienced professionals committed to excellence and continuous improvement, making us a trusted partner for clients across various industries.

About The Role

We are seeking a highly skilled ML Platform Engineer to join our dynamic team. In this role, you will be responsible for designing, building, and maintaining high-performance inference platforms capable of supporting large-scale machine learning models in production environments. Your expertise will ensure that our AI deployment systems are reliable, scalable, and optimized for performance, supporting diverse workloads such as large language models, vision models, and recommendation systems. You will work closely with ML scientists, product teams, and infrastructure engineers to implement robust serving architectures, optimize resource utilization, and enhance system observability. This is an excellent opportunity for professionals passionate about AI infrastructure, performance engineering, and cloud-native deployments to make a significant impact within a forward-thinking organization.

Qualifications

  • Bachelor's or Master's degree in Computer Science, Software Engineering, or a related field
  • Six or more years of experience in distributed systems, infrastructure, or ML platform engineering
  • Proficiency in Python and systems programming languages such as Go, Rust, or C++
  • Hands-on experience operating high-throughput, low-latency services in production environments
  • Experience with large model inference frameworks such as vLLM, TensorRT-LLM, or similar
  • Strong understanding of GPU architecture, memory hierarchies, and accelerator utilization
  • Familiarity with Kubernetes, cloud platforms, and autoscaling mechanisms
  • Experience with observability tools including metrics, tracing, and structured logging
  • Solid background in performance engineering and capacity planning
  • Excellent communication skills and incident response capabilities

Responsibilities

  • Design and operate model serving platforms supporting diverse workloads including LLMs, vision models, and recommendation systems
  • Optimize inference performance through techniques such as continuous batching, request multiplexing, and speculative decoding
  • Implement multi-tenant routing, rate limiting, and quality-of-service policies across model endpoints
  • Build and maintain autoscaling and capacity management systems to balance latency, throughput, and cost
  • Tune GPU utilization, memory management, and cache strategies for large model inference workloads
  • Integrate model serving systems with API gateways, identity management, and observability platforms
  • Implement caching, prompt deduplication, and response reuse strategies to improve efficiency
  • Drive end-to-end observability including latency metrics, queue dynamics, GPU utilization, and error tracking
  • Develop deployment workflows such as canary releases, shadow testing, and automated rollbacks
  • Operate incident response processes for high-availability AI services, ensuring reliability and durability
  • Collaborate with ML and product teams to support new model releases and feature rollouts
  • Implement security controls, including request signing, content filtering, and abuse detection at the serving layer
  • Document operational procedures, performance tuning guides, and system characteristics for internal teams
  • Stay abreast of advancements in AI model serving research and incorporate relevant innovations into production systems

Benefits

  • Competitive salary commensurate with experience
  • 100% remote work within the Continental United States
  • Comprehensive health, dental, and vision insurance plans
  • Paid time off and holidays
  • Opportunities for professional development and continuous learning
  • Long-term, stable employment with a growing company
  • Collaborative and innovative work environment
  • Supportive company culture emphasizing diversity and inclusion

Equal Opportunity

Bright Vision Technologies (BV Teck) is committed to providing equal employment opportunities to all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected characteristic under applicable law. We promote a diverse and inclusive workplace where everyone is valued and respected. We do not tolerate workplace harassment or discrimination and ensure all employment practices are fair and equitable. Our commitment extends to recruitment, hiring, training, promotion, and all other employment-related activities, fostering an environment where all individuals can thrive and succeed.