Create an account for powerful AI tools, award-winning courses, and access to our vibrant community.
Already have an account?
Join 250,000+ professionals and teams at Microsoft, Shopify, and even NASA. 🚀
Already have an account? Login
Find the best remote jobs. Answer a few questions and we'll deploy a powerful assistant to help you search, create alerts, and more.
1 What roles are you open to?
2 Experience level
3 Work style
Did you know? If memory is enabled, Writing.io can remember your job search preferences and help you to improve your resume, craft customized outreach and more.
Category
Leads product strategy and roadmap for Supabase's infrastructure capabilities including database, auth, storage, and edge functions.
Lead technical support for an AI ad management platform, owning customer onboarding, ticket resolution, QA releases, and engineering triage for media buying teams.
Headquarters: Tampa, Florida
URL: https://koast.ai
Head of Technical Support | Full-time | Remote (US hours) | Reports to the CEO
Koast is the AI operating layer for teams that run Meta ads at volume. Agencies, in-house media buying teams, and enterprise buyers managing 10+ ad accounts and launching 50+ ads a week use Koast to launch faster, automate optimizations off real attribution data, and see every account from one Command Center. The Koast Agent is the next layer: an AI that reads the account, spots what matters, and acts inside guardrails.
Our customers are media buyers. When something breaks, it is not a support ticket to them. It is a campaign that did not launch, a budget that did not move, or an account they cannot see. They need someone on the other end who has run ads, knows what they are looking at, and can tell the difference between a Meta API error, a user mistake, and a real bug.
That is this role. You are the anchor between our customers and our engineers. You will onboard every new account, own every ticket, QA every release before it hits production, and make sure engineering hears about patterns, not anecdotes.
What You WIll Own...
Onboarding. Every new customer gets set up right the first time: ad accounts connected, attribution (Hyros, Cometly) wired in, team invited, first launch done together on a call. You will run these sessions yourself, then build the playbook and the in-app flow so the next hundred customers need less of you.
Support, end to end. Intercom is yours. Response times, resolution quality, the help center, the macros, the escalation path. You will answer tickets yourself until the volume justifies a hire, and you will know every customer by name before you delegate a single one.
Triage and the bridge to engineering. You decide what is a bug, what is a feature gap, what is Meta being Meta, and what is user error. Bugs get reproduced, documented with steps and account context, and filed in Shortcut with enough detail that an engineer can fix it without asking you a follow-up question. Patterns get surfaced in the weekly review with numbers: how many customers, how much spend affected, how often.
QA and release quality. Nothing ships to production without you having run it against a real ad account. You will own the QA pass on every release: launch flows, automations, Agent actions, integrations. You know what a broken campaign structure looks like in Ads Manager, and you catch it before a customer does.
System tightness. Meta API connections, token refreshes, webhook health, Whop account provisioning, attribution syncs. You will monitor them (Sentry, PostHog, internal dashboards), notice when something drifts, and get it fixed before it turns into a ticket.
Customer health and retention. You will know which accounts are launching, which have gone quiet, and which are about to churn. You will bring that list to the CEO weekly and you will own the outreach.
Small team, high autonomy. Decisions are made with numbers, written down, and revisited when the numbers change. These are the values we run on. If they sound like you, you will fit right in.
Record a Loom, five minutes or less, covering these four things in order:
Email the Loom link and your LinkedIn or resume to jobs@koast.ai with the subject line "Retention is king".
NOTE: DO NOT APPLY IF YOU HAVE NO EXPERIENCE WITH PRODUCTS LIKE THIS / ADVERTISING / CANNOT WORK US HOURS.
No cover letters. We watch the Loom first and read the resume second. Applications without the video are not reviewed.
To apply: https://weworkremotely.com/remote-jobs/koast-ai-head-of-technical-support-koast-ai
Leads and manages a field sales team to drive revenue growth and client acquisition in the Munich region.
Leads product strategy and development for anti-money laundering and digital asset compliance solutions in a blockchain security company.
Directs product strategy and roadmap for AML and digital asset compliance solutions within a Web3 security platform.
Leads software engineering teams and strategy for travel experiences platform, overseeing technical architecture and engineering excellence.
Directs anti-money laundering and virtual asset compliance programs, ensuring regulatory adherence across blockchain operations.
Leads and manages a field sales team in Munich, driving revenue through direct sales strategy and team performance.
Leads and scales a product support organization, defining support strategy, building technical support teams, and ensuring customer experience aligns with product excellence.
Headquarters: USA (Remote)
Mechanical Orchard is reinventing how the world’s most critical software gets modernized. We’re an applied AI company focused on one of the hardest problems in enterprise technology: rewriting complex legacy systems in a way that is provably correct, low risk, and fast enough to matter. By focusing on system behavior rather than code alone, we turn modernization from a high-stakes, failure-prone effort into a repeatable, confidence-building process that unlocks ongoing innovation.To apply: https://weworkremotely.com/remote-jobs/mechanical-orchard-head-of-product-support
Principal Quality Engineer owns quality strategy and standards, leads the QA squad, and evolves testing tooling and processes across the engineering organization.
Headquarters: Brazil
URL: http://lawnstarter.com
About LawnStarter
LawnStarter is the nation's leading on-demand marketplace for lawn care and related services, with over $150M in annual bookings. We're expanding beyond lawn care to become the one-stop shop for all home services.
About Engineering Quality at LawnStarter
Our QA team is growing, and we're investing real money in better test coverage. We already have shared quality standards and a QA squad that organizes itself horizontally across teams. What we don't have is someone dedicated to owning that vision: every QA is heads-down on their own team, so nobody has the room to drive the standards forward, unify how the squad works, or push the state of the art. That's the gap this role fills. It's not a nice-to-have we're adding because things are going well. It's the structure and leadership we need to make what we've already built actually compound.
The Role
You'll be LawnStarter's first Principal Quality Engineer, reporting directly to the Head of Engineering. We already have quality standards, working CI/CD gates, and AI tooling in the mix. Nobody owns the vision behind all of it, drives it forward, or leads the QA squad that keeps it running. You will.
This isn't a QA management job today. You won't spend your day approving tickets or chasing bug counts. You'll take ownership of the standards and tooling that already exist, work collaboratively with the QA squad to improve them, and keep pushing the state of quality engineering forward as the org grows. As the QA squad grows, this role could take on direct people management of QAs, so hiring, coaching, and performance management experience is a real plus even though it's not the day-one job.
You're not starting from zero, and you're not inheriting a mess either. That means the job isn't "invent a strategy nobody has thought about." It's "listen to the QAs who live in this every day, understand where the pain is, and lead them toward what's next." We need someone who reads where LawnStarter is today and builds from there, not someone who shows up with a favorite tool stack and a coverage number they picked before they met the team.
What makes this role different:
What You'll Own
Problems to Solve
A solid quality bar with no one steering it.
We already have shared standards, working quality gates, and AI tooling in the mix. What's missing is someone whose job is to look at all of it, decide what needs to improve next, and actually drive that improvement instead of it being everyone's part-time responsibility.
A QA squad with no dedicated leadership.
Our QAs already organize horizontally across teams, but every one of them is focused on their own team's work day to day. Nobody has the bandwidth to unify how they work, spread what's working on one team to the others, or represent their pain points where decisions get made. You'll give that squad the structure and leadership it's missing.
Standards that need continuous R&D, not a rewrite.
The foundation works. The risk is standing still while the org grows. You'll keep pushing the state of the art, testing new approaches and tools, and deciding what's worth rolling out broadly versus what stays an experiment.
Quality work that's still someone's side project.
Right now, improving quality practice competes with each QA's day-to-day team commitments. You'll make it someone's actual job to carry that forward, and get engineers, EMs, and PMs treating it as planned work instead of something squeezed in.
What Success Looks Like (Year 1)
Requirements
Who You Are
AI-native, specifically for quality work. You already use AI to generate tests, not just to autocomplete code. You've set up agentic tooling like Claude Code or MCP-based workflows so engineers can run on-demand test generation or code review against their own repos, and you keep experimenting with what AI can take off a team's plate next. This is unlikely to be a good fit if AI tools mostly make you nervous, or your use of them stops at your own personal productivity.
A strategic thinker, not a tool zealot. Give you a team's current maturity level, and you'll design the testing approach that fits where they are, not the one you used at your last job. You resist the urge to mandate a specific framework or a specific coverage number before you understand the context. Someone who shows up insisting "we need to hit 90% coverage" or "everyone must use Playwright" without first learning how the team works will struggle here.
Deeply hands-on across the whole test pyramid. You've built and maintained component, contract, API, E2E, and performance test suites yourself, not just reviewed slides about them. You're the person who steps in when the engineering team is blocked by the pipeline and can't ship. If your test automation experience stops at reading dashboards other people built, this role will feel out of reach fast.
A patient teacher who still ships. Coaching engineers on testing practice and running enablement sessions is part of the job. So is personally building the CI pipelines and the tooling. You don't write the doc and hope someone else implements it. This is unlikely to be a good fit for someone who only wants to advise, or who avoids hands-on infrastructure work.
Comfortable owning influence without owning headcount, at least for now. You'll mentor QA engineers across every squad on technical craft and send notes to their EMs for performance reviews. You won't be their manager on day one, and you won't set their career path. This works well if you're motivated by discipline-wide technical influence. It won't if you need direct reports from the start to feel engaged. If you've hired, coached, and managed performance for a QA or engineering team before, that's a genuine plus: this role could grow into managing the QA squad directly as it scales.
This Role Is NOT
Benefits
Compensation & Benefits
LawnStarter provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability, or genetics. We comply with applicable state and local laws governing nondiscrimination in employment.
To apply: https://weworkremotely.com/remote-jobs/lawnstarter-principal-quality-engineer
Principal Quality Engineer owns quality standards, leads the QA squad, drives testing strategy and tooling evolution, and scales quality practices across the organization.
Headquarters: Mexico
URL: http://lawnstarter.com
About LawnStarter
LawnStarter is the nation's leading on-demand marketplace for lawn care and related services, with over $150M in annual bookings. We're expanding beyond lawn care to become the one-stop shop for all home services.
About Engineering Quality at LawnStarter
Our QA team is growing, and we're investing real money in better test coverage. We already have shared quality standards and a QA squad that organizes itself horizontally across teams. What we don't have is someone dedicated to owning that vision: every QA is heads-down on their own team, so nobody has the room to drive the standards forward, unify how the squad works, or push the state of the art. That's the gap this role fills. It's not a nice-to-have we're adding because things are going well. It's the structure and leadership we need to make what we've already built actually compound.
The Role
You'll be LawnStarter's first Principal Quality Engineer, reporting directly to the Head of Engineering. We already have quality standards, working CI/CD gates, and AI tooling in the mix. Nobody owns the vision behind all of it, drives it forward, or leads the QA squad that keeps it running. You will.
This isn't a QA management job today. You won't spend your day approving tickets or chasing bug counts. You'll take ownership of the standards and tooling that already exist, work collaboratively with the QA squad to improve them, and keep pushing the state of quality engineering forward as the org grows. As the QA squad grows, this role could take on direct people management of QAs, so hiring, coaching, and performance management experience is a real plus even though it's not the day-one job.
You're not starting from zero, and you're not inheriting a mess either. That means the job isn't "invent a strategy nobody has thought about." It's "listen to the QAs who live in this every day, understand where the pain is, and lead them toward what's next." We need someone who reads where LawnStarter is today and builds from there, not someone who shows up with a favorite tool stack and a coverage number they picked before they met the team.
What makes this role different:
What You'll Own
Problems to Solve
A solid quality bar with no one steering it.
We already have shared standards, working quality gates, and AI tooling in the mix. What's missing is someone whose job is to look at all of it, decide what needs to improve next, and actually drive that improvement instead of it being everyone's part-time responsibility.
A QA squad with no dedicated leadership.
Our QAs already organize horizontally across teams, but every one of them is focused on their own team's work day to day. Nobody has the bandwidth to unify how they work, spread what's working on one team to the others, or represent their pain points where decisions get made. You'll give that squad the structure and leadership it's missing.
Standards that need continuous R&D, not a rewrite.
The foundation works. The risk is standing still while the org grows. You'll keep pushing the state of the art, testing new approaches and tools, and deciding what's worth rolling out broadly versus what stays an experiment.
Quality work that's still someone's side project.
Right now, improving quality practice competes with each QA's day-to-day team commitments. You'll make it someone's actual job to carry that forward, and get engineers, EMs, and PMs treating it as planned work instead of something squeezed in.
What Success Looks Like (Year 1)
Requirements
Who You Are
AI-native, specifically for quality work. You already use AI to generate tests, not just to autocomplete code. You've set up agentic tooling like Claude Code or MCP-based workflows so engineers can run on-demand test generation or code review against their own repos, and you keep experimenting with what AI can take off a team's plate next. This is unlikely to be a good fit if AI tools mostly make you nervous, or your use of them stops at your own personal productivity.
A strategic thinker, not a tool zealot. Give you a team's current maturity level, and you'll design the testing approach that fits where they are, not the one you used at your last job. You resist the urge to mandate a specific framework or a specific coverage number before you understand the context. Someone who shows up insisting "we need to hit 90% coverage" or "everyone must use Playwright" without first learning how the team works will struggle here.
Deeply hands-on across the whole test pyramid. You've built and maintained component, contract, API, E2E, and performance test suites yourself, not just reviewed slides about them. You're the person who steps in when the engineering team is blocked by the pipeline and can't ship. If your test automation experience stops at reading dashboards other people built, this role will feel out of reach fast.
A patient teacher who still ships. Coaching engineers on testing practice and running enablement sessions is part of the job. So is personally building the CI pipelines and the tooling. You don't write the doc and hope someone else implements it. This is unlikely to be a good fit for someone who only wants to advise, or who avoids hands-on infrastructure work.
Comfortable owning influence without owning headcount, at least for now. You'll mentor QA engineers across every squad on technical craft and send notes to their EMs for performance reviews. You won't be their manager on day one, and you won't set their career path. This works well if you're motivated by discipline-wide technical influence. It won't if you need direct reports from the start to feel engaged. If you've hired, coached, and managed performance for a QA or engineering team before, that's a genuine plus: this role could grow into managing the QA squad directly as it scales.
This Role Is NOT
Benefits
Compensation & Benefits
LawnStarter provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability, or genetics. We comply with applicable state and local laws governing nondiscrimination in employment.
To apply: https://weworkremotely.com/remote-jobs/lawnstarter-principal-quality-engineer-1
Manages engineering team, oversees project delivery, and leads technical staff across distributed offices and remote locations.
Leads compliance operations and financial crimes prevention programs, ensuring regulatory adherence and risk mitigation across the organization.
Develops and implements AI governance frameworks, policies, and compliance strategies across the organization while partnering with legal, security, and regulatory teams.
Job Summary
The Director of AI Governance sits within Quality, Regulatory, and Compliance and is responsible for designing, developing, and managing DeepHealth’s AI Governance program and strategies. The role promotes safe, legal, and ethical use of AI across the organization in partnership with Information Security, Data Privacy, Vendor Management, Ethics, Legal, and IT.
Essential Duties and Responsibilities
Develop and implement AI Governance strategies aligned with organizational goals and ethical standards.
Design, implement, and maintain an AI Governance framework aligned with applicable standards and regulations, including NIST AI RMF, ISO 42001, and the EU AI Act.
Develop and maintain policies, guidelines, and documentation for responsible development and deployment of AI tools.
Integrate AI governance into existing compliance and risk management processes with cross-functional partners.
Provide expert advice to internal teams and, as needed, to clients on AI governance and compliance matters.
Evaluate vendor compliance with legal, ethical, and security standards related to AI tools and features.
Oversee AI risk assessments and maintain AI inventories and catalogues consistent with legal and regulatory requirements.
Review contracts and draft language related to AI and new technologies in partnership with Legal.
Design and deliver training to build awareness of AI governance principles.
Partner on audits of compliance with the AI framework and applicable legislation, and mentor team members as the function grows.
PLEASE NOTE: This is not an exhaustive list of all duties, responsibilities and requirements of the position described above. Other functions may be assigned and management retains the right to add or change duties at any time.
Minimum Qualifications, Education and Experience
Bachelor’s degree in AI, Cybersecurity, Information Security, or a related field, or equivalent practical experience (required).
5 to 10 years of progressive experience in AI compliance, AI governance, or closely related compliance/risk work in healthcare, medical technology, or another highly regulated industry (required).
Proven experience developing and implementing AI compliance or governance strategies (required).
Working knowledge of AI governance frameworks and regulations such as NIST AI RMF, ISO 42001, and the EU AI Act (required).
Excellent written and oral communication skills, including explaining complex concepts to diverse audiences (required).
High level of integrity and confidentiality (required).
Preferred: Master’s degree or advanced certifications (for example, AAISM, CISSP, CISM) and completed AI governance coursework.
Quality Standards
Communicates, cooperates, and consistently functions professionally and harmoniously with all levels of supervision, co-workers, visitors, and vendors.
Demonstrates initiative, personal awareness, professionalism and integrity, and exercises confidentiality in all areas of performance.
Follows all local, regional and country laws concerning employment.
Follows all DeepHealth policies and procedures.
Follows data privacy, compliance, safety and confidentiality standards at all times.
Practices universal safety precautions.
Promotes good public relations on the phone and in person.
Adapts and is willing to learn new tasks, methods, and systems.
Reports to work regularly as scheduled; consistently punctual with respect to working hours, meal and rest breaks, and maintains satisfactory personal attendance in accordance with DeepHealth guidelines.
Completes job responsibilities in a quality and timely manner.
Travel
This position requires travel up to approximately 10%.
Working Environment
European Union. Remote-friendly or office-based role.
Physical Demands
The employee must be able to perform the essential duties and responsibilities of the position, with or without reasonable accommodation.
Salary:
115,000.00-130,000.00 EURO
100,000.00-110,000,00 France
Site Reliability Engineer Tech Lead builds and maintains critical infrastructure systems, ensures application reliability and SLAs, and leads reliability projects across engineering squads.
ARE YOU INTERESTED IN JOINING AN INNOVATIVE LOGISTICS TECHNOLOGY COMPANY?
Loadsmart is a growth-stage technology company valued at over $1 billion (a true Tech Unicorn)!
We are a collection of industry veterans and user-centered engineers using innovative technology to fearlessly reinvent the future of freight by helping shippers, brokers, warehouses and carriers to move more with less.
With headquarters in Chicago and a globally distributed remote team, Loadsmart continues to attract top talent committed to driving meaningful change. We seek professionals who embody our core values: curiosity, clarity, results, commitment, and teamwork.
In the SRE, Tech Lead role you will build and maintain the company’s internal platform, driving operational excellence and empowering the entire engineering team. You should have experience in analyzing, proposing, and implementing safer systems and processes. Collaborating closely with engineering squads across platform engineering, you will ensure our applications are both safe and reliable. As a hands-on leader, you will stay actively involved in technical work while collaborating closely with internal stakeholders and engineering squads across the organization to ensure our applications are both safe and reliable.
DEPARTMENT:Â Engineering
LOCATION: Anywhere in Brazil - Remote
WORKING AT LOADSMART:
• Competitive base salaries - we believe in rewarding top talent
• Extremely competitive Equity package - become a shareholder in our company!
• Loadie Time Off - PTO and sick days without a limit
At Loadsmart, we believe our biggest asset is our people. We are proud to be an equal opportunity employer, hiring and developing individuals from diverse backgrounds and experiences to add to our collaborative culture. Loadsmart treats all candidates and employees with respect and does not discriminate in our recruiting, hiring, and promoting processes, including on the basis of race, color, religion, sex, age, sexual orientation, gender identity and/or expression, national origin, veteran status, or disability.
It is the policy of Loadsmart that all offers of employment made shall be contingent upon successful completion of electronic background check(s). These checks will be job-related, consistent with business necessity and conducted by our vendor, pursuant to all applicable laws, rules, policies and procedures of our candidates’ specific locale.
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
Lead GPU infrastructure engineering for a Kubernetes-based compute platform, managing containerized GPU resources and inference endpoints at scale.
Join Tether and Shape the Future of Digital Finance
At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.
Innovate with Tether
Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.
But that’s just the beginning:
Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.
Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.
Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.
Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.
Why Join Us?
Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.
If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.
Are you ready to be part of the future?
About the job
Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.
The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.
This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.
Responsibilities
Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.
Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.
Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.
Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.
Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.
Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.
Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.
Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.
Hiring. Complete the platform team and set the technical bar for the engineers who join it.
Must have
Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.
Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.
GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.
High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.
Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.
Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.
HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.
Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.
Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.
A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.
Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.
Excellent written and spoken English. Most partner and leadership work happens in writing.
Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.
Desirable
Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).
Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.
VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).
Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.
Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.
Peer-to-peer or distributed-systems background.
Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.
Important information for candidates
Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:
Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/
Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.
Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.
Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io
We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.
When in doubt, feel free to reach out through our official website.
Technical Lead oversees GPU infrastructure and Kubernetes platform development for Tether Data's managed inference and compute services.
Join Tether and Shape the Future of Digital Finance
At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.
Innovate with Tether
Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.
But that’s just the beginning:
Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.
Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.
Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.
Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.
Why Join Us?
Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.
If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.
Are you ready to be part of the future?
About the job
Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.
The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.
This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.
Responsibilities
Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.
Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.
Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.
Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.
Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.
Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.
Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.
Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.
Hiring. Complete the platform team and set the technical bar for the engineers who join it.
Must have
Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.
Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.
GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.
High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.
Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.
Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.
HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.
Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.
Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.
A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.
Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.
Excellent written and spoken English. Most partner and leadership work happens in writing.
Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.
Desirable
Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).
Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.
VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).
Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.
Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.
Peer-to-peer or distributed-systems background.
Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.
Important information for candidates
Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:
Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/
Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.
Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.
Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io
We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.
When in doubt, feel free to reach out through our official website.
Technical Lead manages GPU infrastructure and Kubernetes-based compute platform, overseeing architecture, performance, and team delivery for a managed inference service.
Join Tether and Shape the Future of Digital Finance
At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.
Innovate with Tether
Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.
But that’s just the beginning:
Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.
Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.
Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.
Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.
Why Join Us?
Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.
If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.
Are you ready to be part of the future?
About the job
Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.
The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.
This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.
Responsibilities
Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.
Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.
Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.
Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.
Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.
Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.
Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.
Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.
Hiring. Complete the platform team and set the technical bar for the engineers who join it.
Must have
Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.
Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.
GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.
High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.
Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.
Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.
HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.
Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.
Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.
A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.
Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.
Excellent written and spoken English. Most partner and leadership work happens in writing.
Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.
Desirable
Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).
Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.
VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).
Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.
Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.
Peer-to-peer or distributed-systems background.
Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.
Important information for candidates
Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:
Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/
Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.
Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.
Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io
We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.
When in doubt, feel free to reach out through our official website.
Designs and manages AI agent infrastructure on Amazon Bedrock, including orchestration, retrieval pipelines, security controls, and tools for team-wide AI feature deployment.
Bedrock Ocean builds and operates autonomous underwater vehicles (AUVs) that collect georeferenced ocean-floor data at commercial scale. We deliver bathymetric and imagery data products to customers through our own platform, and we’re scaling toward continuous, around-the-clock data collection campaigns spanning months at a time.
We are building AI agents on Amazon Bedrock to support our ocean data, internal operations, and customer platform. This role owns that architecture.
(One note on names. Amazon Bedrock is the AWS service. Bedrock Ocean is us. They are unrelated, and we are aware it is confusing.)
We are looking for a Staff Platform Engineer to lead our AI architecture. This role goes beyond building agents on existing platforms; you will create the infrastructure itself, including the orchestration layer, the data and retrieval pipeline, and the security model required to work with production data. You will also build the tools and abstractions that allow our engineering team to implement AI features independently.
This position combines software engineering, data engineering, and infrastructure operations. You will manage the full lifecycle of our Amazon Bedrock implementation, from initial data chunking to IAM access controls. While some of our data pipelines are already in place, they will require significant expansion, and others will need to be built from scratch.
Security is core to this role, not an afterthought. Because agents with tool access represent a new kind of system actor, you will define their operational boundaries, including what they can access, the actions they can perform autonomously, and the monitoring required to detect issues.
Our roadmap prioritizes internal engineering and operational systems first to ensure a fast feedback loop, followed by our ocean and survey data products. Customer-facing retrieval is the final, high-stakes phase. You will play a key role in defining this sequence. While you will not be responsible for the core data transport design (store-and-forward or hub-and-spoke), you will work closely with that team to ensure it meets our retrieval and data freshness requirements.
Architect Agent Orchestration: Design the Amazon Bedrock integration, including agent and action group configuration, backend APIs, model access, throughput, and cross-environment deployment.
Manage Retrieval Data Plane: Own the end-to-end retrieval pipeline from ingestion and chunking to embedding and storage in Amazon OpenSearch Serverless. Focus on optimizing for index design, cost, and capacity.
Extend Data Pipelines: Adapt ingestion pipelines for internal knowledge, ocean data, and customer platforms, addressing challenges specific to geospatial and large-binary datasets.
Secure AI Infrastructure: Implement robust security including Bedrock Guardrails, VPC and PrivateLink network boundaries, least-privilege IAM, and audit trails to ensure data isolation.
Define Agent Governance: Build the mechanisms to enforce approval boundaries for autonomous actions, ensuring agents are safe and monitored.
Establish LLMOps & Observability: Implement comprehensive monitoring for tracing, tool calls, and retrieval performance, using CloudWatch and LLM-specific tools like Langfuse or Phoenix.
Build Evaluation Frameworks: Create the infrastructure to run automated evaluations, track results, and manage release gates for model accuracy.
Enable Engineering Productivity: Provide the team with abstraction layers, SDKs, and self-service environments that allow engineers to ship AI features independently.
Operational Excellence: Manage the environment as code across all stages, ensuring deployment safety and participating in incident reviews.
8+ years in software and infrastructure engineering, including deep production backend experience (Python or TypeScript preferred, Go fine) and staff-level ownership of technical direction.
Hands-on experience standing up Amazon Bedrock in production: agents, knowledge bases, guardrails, model access, and the throughput and quota decisions that come with them.
Containerized service deployment on ECS, EKS, or Lambda, with CI/CD you have owned rather than inherited. The models are managed, but the backend APIs, tool endpoints, and ingestion jobs still run somewhere real.
Practical RAG and vector search experience: embeddings, chunking strategies, semantic search quality, and operating a managed vector database (OpenSearch Serverless, Pinecone, pgvector, or similar) at production scale and cost.
Real data engineering: you have built or substantially extended ingestion pipelines over messy, heterogeneous, unstructured sources, and you think about freshness and correctness as SLAs rather than afterthoughts.
Strong AWS ecosystem expertise: IAM roles and least privilege for machine identities, VPC networking and PrivateLink, Lambda, S3, KMS, CloudWatch, and provisioning safely through infrastructure as code (Terraform, CDK, or CloudFormation).
Production LLM exposure: you have moved LLM features or autonomous agents past the prototype stage into environments other people depend on.
A working point of view on securing agentic systems: scoping tool permissions, prompt injection and exfiltration risk, sensitive data handling in retrieval, and where a human belongs in the loop.
Experience designing developer-facing APIs, SDKs, or platform services with an API-first mindset. The interface is the product for the engineers who consume it.
Experience building and operating multi-tenant services, with isolation guarantees that hold when the data belongs to customers rather than to us.
Platform instinct: you build the abstraction other engineers stand on, and you measure yourself by what they ship rather than by what you ship directly.
Demonstrated technical leadership and system design judgment at staff level: you have driven an architectural direction across pods or teams you do not manage, and made it stick through influence rather than authority.
A pragmatic builder’s bias. You reach for boring, fully managed infrastructure before complex self-hosted alternatives, and you can tell the difference between the two in an architecture review.
Comfort wearing several hats on a small team, and the discipline to write things down so the system runs without you.
Experience deploying LLM evaluations to measure accuracy over time and treating eval results as a release gate.
Involvement in AI red-teaming or the AI security community.
Experience with GraphRAG or knowledge graphs.
Experience running retrieval over geospatial, scientific, or large-binary datasets.
Experience moving data across intermittent or unreliable links: store-and-forward, hub-and-spoke topologies, offload from disconnected or edge systems, and reconciliation once a link comes back.
Compliance experience such as SOC 2, or handling government or defense customer data.
Background supporting data platforms, autonomous systems, or field operations.
Your AI work has been prototypes and notebooks rather than systems other people depend on in production.
You want to be a model researcher, a prompt engineer, or to spend your time fine-tuning models. This role owns the platform underneath agents and partners closely with the people building them.
You treat security as a gate at the end of a project rather than something designed from the start.
You would rather self-host and build from scratch than adopt a managed service that already works. We are optimizing for a small team shipping, not for architectural purity.
You want a mature platform team and a narrow, well-bounded scope. This is an early build with a lot of surface area and few existing answers.
At Bedrock Ocean, our mission is to make the ocean transparent. We are building more than just a survey service; we are creating a source of deep ocean intelligence that grows with every mission.
To realize this, we need to make our data accessible and actionable. This role is about building the infrastructure that allows our teams and eventually our customers to directly query and learn from our findings. Because this data is strategically sensitive, security is woven into the foundation of your work, not added on later.
Ultimately, an AI agent with access to our production data is a powerful tool, but it requires careful design to be a success. You will build a secure, reliable platform that ensures these systems empower our mission safely and effectively.
The base compensation for this role is expected to be $160,000- $220,000 annually plus equity.
Bedrock Ocean is an equal opportunity employer.
Senior leader accountable for pod delivery, margin performance, and cross-functional execution across a portfolio of AI data and engagement work.
About Snorkel
At Snorkel, we believe meaningful AI doesn’t start with the model, it starts with the data.
We’re on a mission to help enterprises transform expert knowledge into specialized AI at scale. The AI landscape has gone through incredible changes since 2015, when Snorkel started as a research project in the Stanford AI Lab, to the generative AI breakthroughs of today. But one thing has remained constant: the data you use to build AI is the key to achieving differentiation, high performance, and production-ready systems. We work with some of the world’s largest organizations to empower scientists, engineers, financial experts, product creators, journalists, and more to build custom AI with their data faster than ever before. Excited to help us redefine how AI is built? Apply to be the newest Snorkeler!
The Strategic Delivery Lead is the senior operator accountable for turning a pod’s ambition into reliably delivered, high-quality work. This role is hybrid ( _3 days/week in office)_ in San Francisco,CA or New York City, NY. You will own the performance of a portfolio spanning samples, off-the-shelf datasets, proprietary data products, and bespoke engagements, connecting GTM, Research, Delivery, Product, Engineering, and Expert AI Supply to ensure the right work is prioritized, staffed, and executed. This is a highly cross-functional role: you will lead through influence cross functionally and directly manage a team of strategic project leads (SPLs), bring clarity to complex tradeoffs, and personally intervene when quality, delivery, or customer outcomes are at risk.
You will be the pod’s owner for margin performance and operating economics, with a clear view of cost-to-deliver, gross margin, capacity, and the operational choices that drive them. While you will not set pricing, your cost, scope, and delivery inputs will directly inform pricing and commercial decisions. In partnership with Product, Research and Forward Deployed Engineering, you will translate the domain’s strategy and priority bets into an executable roadmap, mobilizing the teams required to bring it to life and building the systems, rhythms, and playbooks that make the model stronger with every launch.
Bonus:Â experience in AI/ML data, evals or benchmarks; having built a function or business line inside a fast-scaling company; a founder or early-employee background in a technical product.
Pay Transparency Notice:Â Depending on your work location, the target annual salary base for this position can range as detailed below. Snorkel also includes benefits (including medical, dental, vision and 401(k)). The salary range for this position based off of tier 1 locations such as San Francisco Bay Area, and New York, and is $296,000 - $444,000. All offers include equity compensation.
Be Your Best at Snorkel
Joining Snorkel AI means becoming part of a company that has market proven solutions, robust funding, and is scaling rapidly—offering a unique combination of stability and the excitement of high growth. As a member of our team, you’ll have meaningful opportunities to shape priorities and initiatives, influence key strategic decisions, and directly impact our ongoing success. Whether you’re looking to deepen your technical expertise, explore leadership opportunities, or learn new skills across multiple functions, you’re fully supported in building your career in an environment designed for growth, learning, and shared success.
Snorkel AI is proud to be an Equal Employment Opportunity employer and is committed to building a team that represents a variety of backgrounds, perspectives, and skills. Snorkel AI embraces diversity and provides equal employment opportunities to all employees and applicants for employment. Snorkel AI prohibits discrimination and harassment of any type on the basis of race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state, or local law. All employment is decided on the basis of qualifications, performance, merit, and business need.
We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us to request accommodation.