Create an account for powerful AI tools, award-winning courses, and access to our vibrant community.
Already have an account?
Join 250,000+ professionals and teams at Microsoft, Shopify, and even NASA. 🚀
Already have an account? Login
Find the best remote jobs. Answer a few questions and we'll deploy a powerful assistant to help you search, create alerts, and more.
1 What roles are you open to?
2 Experience level
3 Work style
Did you know? If memory is enabled, Writing.io can remember your job search preferences and help you to improve your resume, craft customized outreach and more.
Category
Lead technical support for an AI ad management platform, owning customer onboarding, ticket resolution, QA releases, and engineering triage for media buying teams.
Headquarters: Tampa, Florida
URL: https://koast.ai
Head of Technical Support | Full-time | Remote (US hours) | Reports to the CEO
Koast is the AI operating layer for teams that run Meta ads at volume. Agencies, in-house media buying teams, and enterprise buyers managing 10+ ad accounts and launching 50+ ads a week use Koast to launch faster, automate optimizations off real attribution data, and see every account from one Command Center. The Koast Agent is the next layer: an AI that reads the account, spots what matters, and acts inside guardrails.
Our customers are media buyers. When something breaks, it is not a support ticket to them. It is a campaign that did not launch, a budget that did not move, or an account they cannot see. They need someone on the other end who has run ads, knows what they are looking at, and can tell the difference between a Meta API error, a user mistake, and a real bug.
That is this role. You are the anchor between our customers and our engineers. You will onboard every new account, own every ticket, QA every release before it hits production, and make sure engineering hears about patterns, not anecdotes.
What You WIll Own...
Onboarding. Every new customer gets set up right the first time: ad accounts connected, attribution (Hyros, Cometly) wired in, team invited, first launch done together on a call. You will run these sessions yourself, then build the playbook and the in-app flow so the next hundred customers need less of you.
Support, end to end. Intercom is yours. Response times, resolution quality, the help center, the macros, the escalation path. You will answer tickets yourself until the volume justifies a hire, and you will know every customer by name before you delegate a single one.
Triage and the bridge to engineering. You decide what is a bug, what is a feature gap, what is Meta being Meta, and what is user error. Bugs get reproduced, documented with steps and account context, and filed in Shortcut with enough detail that an engineer can fix it without asking you a follow-up question. Patterns get surfaced in the weekly review with numbers: how many customers, how much spend affected, how often.
QA and release quality. Nothing ships to production without you having run it against a real ad account. You will own the QA pass on every release: launch flows, automations, Agent actions, integrations. You know what a broken campaign structure looks like in Ads Manager, and you catch it before a customer does.
System tightness. Meta API connections, token refreshes, webhook health, Whop account provisioning, attribution syncs. You will monitor them (Sentry, PostHog, internal dashboards), notice when something drifts, and get it fixed before it turns into a ticket.
Customer health and retention. You will know which accounts are launching, which have gone quiet, and which are about to churn. You will bring that list to the CEO weekly and you will own the outreach.
Small team, high autonomy. Decisions are made with numbers, written down, and revisited when the numbers change. These are the values we run on. If they sound like you, you will fit right in.
Record a Loom, five minutes or less, covering these four things in order:
Email the Loom link and your LinkedIn or resume to jobs@koast.ai with the subject line "Retention is king".
NOTE: DO NOT APPLY IF YOU HAVE NO EXPERIENCE WITH PRODUCTS LIKE THIS / ADVERTISING / CANNOT WORK US HOURS.
No cover letters. We watch the Loom first and read the resume second. Applications without the video are not reviewed.
To apply: https://weworkremotely.com/remote-jobs/koast-ai-head-of-technical-support-koast-ai
Leads and scales a product support organization, defining support strategy, building technical support teams, and ensuring customer experience aligns with product excellence.
Headquarters: USA (Remote)
Mechanical Orchard is reinventing how the world’s most critical software gets modernized. We’re an applied AI company focused on one of the hardest problems in enterprise technology: rewriting complex legacy systems in a way that is provably correct, low risk, and fast enough to matter. By focusing on system behavior rather than code alone, we turn modernization from a high-stakes, failure-prone effort into a repeatable, confidence-building process that unlocks ongoing innovation.To apply: https://weworkremotely.com/remote-jobs/mechanical-orchard-head-of-product-support
Manages engineering team, oversees project delivery, and leads technical staff across distributed offices and remote locations.
Site Reliability Engineer Tech Lead builds and maintains critical infrastructure systems, ensures application reliability and SLAs, and leads reliability projects across engineering squads.
ARE YOU INTERESTED IN JOINING AN INNOVATIVE LOGISTICS TECHNOLOGY COMPANY?
Loadsmart is a growth-stage technology company valued at over $1 billion (a true Tech Unicorn)!
We are a collection of industry veterans and user-centered engineers using innovative technology to fearlessly reinvent the future of freight by helping shippers, brokers, warehouses and carriers to move more with less.
With headquarters in Chicago and a globally distributed remote team, Loadsmart continues to attract top talent committed to driving meaningful change. We seek professionals who embody our core values: curiosity, clarity, results, commitment, and teamwork.
In the SRE, Tech Lead role you will build and maintain the company’s internal platform, driving operational excellence and empowering the entire engineering team. You should have experience in analyzing, proposing, and implementing safer systems and processes. Collaborating closely with engineering squads across platform engineering, you will ensure our applications are both safe and reliable. As a hands-on leader, you will stay actively involved in technical work while collaborating closely with internal stakeholders and engineering squads across the organization to ensure our applications are both safe and reliable.
DEPARTMENT:Â Engineering
LOCATION: Anywhere in Brazil - Remote
WORKING AT LOADSMART:
• Competitive base salaries - we believe in rewarding top talent
• Extremely competitive Equity package - become a shareholder in our company!
• Loadie Time Off - PTO and sick days without a limit
At Loadsmart, we believe our biggest asset is our people. We are proud to be an equal opportunity employer, hiring and developing individuals from diverse backgrounds and experiences to add to our collaborative culture. Loadsmart treats all candidates and employees with respect and does not discriminate in our recruiting, hiring, and promoting processes, including on the basis of race, color, religion, sex, age, sexual orientation, gender identity and/or expression, national origin, veteran status, or disability.
It is the policy of Loadsmart that all offers of employment made shall be contingent upon successful completion of electronic background check(s). These checks will be job-related, consistent with business necessity and conducted by our vendor, pursuant to all applicable laws, rules, policies and procedures of our candidates’ specific locale.
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
Lead GPU infrastructure engineering for a Kubernetes-based compute platform, managing containerized GPU resources and inference endpoints at scale.
Join Tether and Shape the Future of Digital Finance
At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.
Innovate with Tether
Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.
But that’s just the beginning:
Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.
Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.
Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.
Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.
Why Join Us?
Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.
If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.
Are you ready to be part of the future?
About the job
Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.
The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.
This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.
Responsibilities
Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.
Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.
Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.
Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.
Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.
Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.
Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.
Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.
Hiring. Complete the platform team and set the technical bar for the engineers who join it.
Must have
Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.
Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.
GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.
High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.
Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.
Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.
HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.
Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.
Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.
A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.
Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.
Excellent written and spoken English. Most partner and leadership work happens in writing.
Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.
Desirable
Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).
Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.
VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).
Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.
Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.
Peer-to-peer or distributed-systems background.
Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.
Important information for candidates
Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:
Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/
Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.
Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.
Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io
We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.
When in doubt, feel free to reach out through our official website.
Technical Lead oversees GPU infrastructure and Kubernetes platform development for Tether Data's managed inference and compute services.
Join Tether and Shape the Future of Digital Finance
At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.
Innovate with Tether
Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.
But that’s just the beginning:
Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.
Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.
Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.
Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.
Why Join Us?
Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.
If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.
Are you ready to be part of the future?
About the job
Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.
The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.
This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.
Responsibilities
Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.
Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.
Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.
Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.
Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.
Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.
Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.
Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.
Hiring. Complete the platform team and set the technical bar for the engineers who join it.
Must have
Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.
Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.
GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.
High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.
Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.
Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.
HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.
Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.
Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.
A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.
Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.
Excellent written and spoken English. Most partner and leadership work happens in writing.
Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.
Desirable
Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).
Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.
VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).
Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.
Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.
Peer-to-peer or distributed-systems background.
Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.
Important information for candidates
Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:
Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/
Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.
Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.
Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io
We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.
When in doubt, feel free to reach out through our official website.
Technical Lead manages GPU infrastructure and Kubernetes-based compute platform, overseeing architecture, performance, and team delivery for a managed inference service.
Join Tether and Shape the Future of Digital Finance
At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.
Innovate with Tether
Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.
But that’s just the beginning:
Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.
Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.
Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.
Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.
Why Join Us?
Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.
If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.
Are you ready to be part of the future?
About the job
Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.
The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.
This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.
Responsibilities
Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.
Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.
Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.
Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.
Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.
Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.
Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.
Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.
Hiring. Complete the platform team and set the technical bar for the engineers who join it.
Must have
Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.
Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.
GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.
High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.
Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.
Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.
HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.
Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.
Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.
A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.
Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.
Excellent written and spoken English. Most partner and leadership work happens in writing.
Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.
Desirable
Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).
Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.
VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).
Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.
Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.
Peer-to-peer or distributed-systems background.
Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.
Important information for candidates
Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:
Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/
Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.
Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.
Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io
We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.
When in doubt, feel free to reach out through our official website.
Leads a team of Account Executives selling enterprise SaaS software productivity solutions to CIOs, CTOs, and engineering leaders across EMEA, managing complex deals and coaching team performance.
BlueOptima gives CIOs, CTOs and CFOs objective, code-level visibility into software productivity, AI ROI, and engineering cost. This is a problem that until recently could not be measured at all. With AI-generated code now flowing into enterprise codebases at scale, that measurement gap has become a board-level priority.
10B+ code revisions analysed · 800,000+ developers on platform · 9 of the top 16 universal banks
The business is profitable, bootstrapped, and 20 years deep in enterprise relationships. You are joining a sales organisation that is mid-build. The foundations are in place, the product has genuine enterprise traction, and there is real room to grow fast.
Buyers are CIOs, CTOs, VP Engineering and CFO stakeholders at large software engineering organisations. The conversation is about developer productivity, engineering cost, and AI ROI. These are budget-backed problems that sit at the top of the technology agenda right now. You are not convincing anyone the problem exists.
You will lead a team of Account Executives working on enterprise accounts across the UK and EMEA, with deal sizes typically in the ÂŁ100k to ÂŁ200k ARR range. You will be cultivating a culture of success and continuous improvement through your coaching. Your leadership will guide the team through complex sales cycles and ambitious growth targets.
You will use sales data and team performance metrics to identify opportunities for improvement. Implement actionable insights to enhance sales tactics and team productivity, leveraging coaching frameworks to work side by side with your team members.
You will work in close partnership with Marketing and SDR leadership to identify new pipeline generation opportunities, and with Customer Success leadership to surface expansion opportunities within the install base.
Required
Valued
Why This Role Makes Sense Right Now
AI governance and engineering ROI measurement have moved from nice-to-have to board priority in the last 18 months. The buyers are funded, the problem is urgent, and BlueOptima is one of the very few companies with a production-grade answer. For a sales leader who wants to work with a team closing real enterprise deals, not chasing SMB volume, the timing is good.
Where This Leads
The natural progression from this role is into second-line leadership. That path is defined and is based on performance and the growth of the company, not tenure.
Benefits we offer
Stay connected with us on LinkedIn or keep an eye on our career page for future opportunities!
Define and implement SLIs/SLOs, establish error budgets, strengthen incident response processes, and coach engineering teams on reliability best practices across the platform.
Fingerprint empowers enterprises to detect and stop online fraud with the world’s most accurate device intelligence. We lead our industry with bleeding-edge identification capabilities and work on turning new ideas and discoveries in the fraud detection space into reality. Our customers range from innovative startups to leading enterprise companies, including Plaid, Dropbox, and Booking.com.
Fingerprint is a globally dispersed, 100% remote company. We were named on on the 2026 Forbes Best Startup Employers list and ranked #803 on the 2026 Inc. 5000 list of America’s fastest-growing private companies.
We have raised $77M and are backed by Craft Ventures (Tesla, Facebook, Airbnb ), Nexus Venture Partners ( Postman, Apollo.io, MinIO, Druva) and Uncorrelated Ventures ( Redis, Rollbar, Â Gradle).
You will be Fingerprint’s first dedicated Site Reliability Engineer. You will work alongside our Architect on the shape of the platform, with Cloud Platform on the infrastructure that runs it, and with every product team on how they operate what they own — without direct reports. The mandate has three parts. First, make reliability measurable: define SLIs and SLOs for our critical paths, get teams to own them, and make error budgets the shared language for prioritizing reliability against features. Second, raise the operational bar: strengthen incident response, postmortem quality, alerting, and change safety so that we find out first and repeat incidents stop repeating. Third, build the mindset: coach teams to design for failure, test for it deliberately, and treat operability as part of done — so the practices outlive your involvement in any single team.
You report directly to the VP of Engineering. That placement is deliberate: reliability standards need to apply evenly across eight engineering groups, and you need the neutrality to hold every team, including infrastructure, to the same bar.
Make reliability measurable
Raise the operational bar
Build the SRE mindset in teams
Stay hands-on
Nice to have
Compensation & Transparency
For US-based employees, the cash compensation range for this role is $177,000 – $240,000. We set standard ranges for all US roles based on function, level, and geographic location, benchmarked against similar stage growth companies. To comply with local legislation and provide greater transparency, we share salary ranges on all job postings. However, these ranges are specific to the hiring location and may differ within or outside the US. Offers vary depending on, but not limited to, relevant experience, education, certifications/licenses, skills, training, and market conditions.
Due to regulatory and security reasons, there’s a small number of countries where we cannot have Fingerprint teammates based. Additionally, because Fingerprint is an all-remote company and people can join our workforce from almost any country, we do not sponsor visas. Fingerprint teammates need to be authorized to work from their home location.
We are dedicated to creating an inclusive work environment for everyone. We embrace and celebrate the unique experiences, perspectives and cultural backgrounds that each employee brings to our workplace. Fingerprint strives to foster an environment where our employees feel respected, valued and empowered, and our team members are at the forefront in helping us promote and sustain an inclusive workplace. We highly encourage people from underrepresented groups in tech to apply.
If you are applying as a resident of California, please read our CCPA notice here.
If you are applying as a resident of the EU, please read our GDPR notice here.
Due to regulatory and security reasons, there’s a small number of countries where we cannot have Fingerprint teammates based. Additionally, because Fingerprint is an all-remote company and people can join our workforce from almost any country, we do not sponsor visas. Fingerprint teammates need to be authorized to work from their home location.
We are dedicated to creating an inclusive work environment for everyone. We embrace and celebrate the unique experiences, perspectives and cultural backgrounds that each employee brings to our workplace. Fingerprint strives to foster an environment where our employees feel respected, valued and empowered, and our team members are at the forefront in helping us promote and sustain an inclusive workplace. We highly encourage people from underrepresented groups in tech to apply.
If you are applying as a resident of California, please read our CCPA notice here.
If you are applying as a resident of the EU, please read our GDPR notice here.
**We have noticed a rise in recruiting impersonations across the industry, where scammers attempt to access candidates’ personal and financial information through fake interviews and offers. All Fingerprint recruiting email communications will always come from the @fingerprint.com domain. Any outreach claiming to be from Fingerprint via other sources should be ignored.
Lead cross-cutting architecture decisions across engineering teams, designing scalable systems and setting technical standards for a fraud detection platform.
Fingerprint empowers enterprises to detect and stop online fraud with the world’s most accurate device intelligence. We lead our industry with bleeding-edge identification capabilities and work on turning new ideas and discoveries in the fraud detection space into reality. Our customers range from innovative startups to leading enterprise companies, including Plaid, Dropbox, and Booking.com.
Fingerprint is a globally dispersed, 100% remote company. We were named on on the 2026 Forbes Best Startup Employers list and ranked #803 on the 2026 Inc. 5000 list of America’s fastest-growing private companies.
We have raised $77M and are backed by Craft Ventures (Tesla, Facebook, Airbnb ), Nexus Venture Partners ( Postman, Apollo.io, MinIO, Druva) and Uncorrelated Ventures ( Redis, Rollbar, Â Gradle).
You will be Fingerprint’s first dedicated Architect. You will lead cross-cutting architecture — how our systems fit together and where they need to go — working alongside the Staff and Lead engineers who already shape them, without direct reports. The mandate has three parts. First, go deep: build the end-to-end picture nobody currently has time to hold, identify where the platform will strain as we scale, and set the direction to address it. Second, raise the bar: strengthen and extend our design review, API, service, and reliability practices so eight engineering groups can move fast without stepping on each other. Third, make it durable: give cross-cutting architecture work consistent cadence, a durable decision record, and follow-through, so strong individual judgment compounds into platform-level outcomes.
You report directly to the VP of Engineering. That placement is deliberate — you need neutrality across groups so the standards you set with teams are adopted everywhere.
Own the end-to-end architecture
Lead cross-cutting architecture
Strengthen and extend best practices
Set us up for scale
Nice to have
Compensation & Transparency
For US-based employees, the cash compensation range for this role is $177,000 – $240,000. We set standard ranges for all US roles based on function, level, and geographic location, benchmarked against similar stage growth companies. To comply with local legislation and provide greater transparency, we share salary ranges on all job postings. However, these ranges are specific to the hiring location and may differ within or outside the US. Offers vary depending on, but not limited to, relevant experience, education, certifications/licenses, skills, training, and market conditions.
Due to regulatory and security reasons, there’s a small number of countries where we cannot have Fingerprint teammates based. Additionally, because Fingerprint is an all-remote company and people can join our workforce from almost any country, we do not sponsor visas. Fingerprint teammates need to be authorized to work from their home location.
We are dedicated to creating an inclusive work environment for everyone. We embrace and celebrate the unique experiences, perspectives and cultural backgrounds that each employee brings to our workplace. Fingerprint strives to foster an environment where our employees feel respected, valued and empowered, and our team members are at the forefront in helping us promote and sustain an inclusive workplace. We highly encourage people from underrepresented groups in tech to apply.
If you are applying as a resident of California, please read our CCPA notice here.
If you are applying as a resident of the EU, please read our GDPR notice here.
Due to regulatory and security reasons, there’s a small number of countries where we cannot have Fingerprint teammates based. Additionally, because Fingerprint is an all-remote company and people can join our workforce from almost any country, we do not sponsor visas. Fingerprint teammates need to be authorized to work from their home location.
We are dedicated to creating an inclusive work environment for everyone. We embrace and celebrate the unique experiences, perspectives and cultural backgrounds that each employee brings to our workplace. Fingerprint strives to foster an environment where our employees feel respected, valued and empowered, and our team members are at the forefront in helping us promote and sustain an inclusive workplace. We highly encourage people from underrepresented groups in tech to apply.
If you are applying as a resident of California, please read our CCPA notice here.
If you are applying as a resident of the EU, please read our GDPR notice here.
**We have noticed a rise in recruiting impersonations across the industry, where scammers attempt to access candidates’ personal and financial information through fake interviews and offers. All Fingerprint recruiting email communications will always come from the @fingerprint.com domain. Any outreach claiming to be from Fingerprint via other sources should be ignored.
Technical Lead oversees GPU infrastructure and Kubernetes-based compute platform development, managing architecture and team for Tether Data's managed inference services.
Join Tether and Shape the Future of Digital Finance
At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.
Innovate with Tether
Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.
But that’s just the beginning:
Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.
Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.
Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.
Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.
Why Join Us?
Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.
If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.
Are you ready to be part of the future?
About the job
Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.
The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.
This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.
Responsibilities
Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.
Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.
Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.
Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.
Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.
Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.
Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.
Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.
Hiring. Complete the platform team and set the technical bar for the engineers who join it.
Must have
Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.
Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.
GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.
High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.
Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.
Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.
HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.
Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.
Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.
A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.
Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.
Excellent written and spoken English. Most partner and leadership work happens in writing.
Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.
Desirable
Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).
Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.
VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).
Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.
Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.
Peer-to-peer or distributed-systems background.
Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.
Important information for candidates
Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:
Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/
Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.
Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.
Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io
We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.
When in doubt, feel free to reach out through our official website.
Technical Lead designs and manages GPU infrastructure and Kubernetes-based compute platform for AI inference and managed services.
Join Tether and Shape the Future of Digital Finance
At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.
Innovate with Tether
Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.
But that’s just the beginning:
Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.
Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.
Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.
Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.
Why Join Us?
Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.
If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.
Are you ready to be part of the future?
About the job
Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.
The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.
This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.
Responsibilities
Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.
Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.
Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.
Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.
Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.
Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.
Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.
Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.
Hiring. Complete the platform team and set the technical bar for the engineers who join it.
Must have
Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.
Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.
GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.
High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.
Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.
Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.
HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.
Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.
Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.
A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.
Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.
Excellent written and spoken English. Most partner and leadership work happens in writing.
Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.
Desirable
Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).
Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.
VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).
Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.
Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.
Peer-to-peer or distributed-systems background.
Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.
Important information for candidates
Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:
Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/
Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.
Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.
Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io
We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.
When in doubt, feel free to reach out through our official website.
Technical Lead oversees GPU infrastructure and Kubernetes-based platform development, managing compute resources and managed inference endpoints for Tether Data's AI infrastructure.
Join Tether and Shape the Future of Digital Finance
At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.
Innovate with Tether
Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.
But that’s just the beginning:
Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.
Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.
Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.
Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.
Why Join Us?
Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.
If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.
Are you ready to be part of the future?
About the job
Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.
The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.
This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.
Responsibilities
Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.
Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.
Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.
Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.
Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.
Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.
Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.
Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.
Hiring. Complete the platform team and set the technical bar for the engineers who join it.
Must have
Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.
Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.
GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.
High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.
Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.
Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.
HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.
Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.
Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.
A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.
Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.
Excellent written and spoken English. Most partner and leadership work happens in writing.
Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.
Desirable
Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).
Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.
VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).
Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.
Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.
Peer-to-peer or distributed-systems background.
Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.
Important information for candidates
Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:
Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/
Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.
Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.
Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io
We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.
When in doubt, feel free to reach out through our official website.
Site Reliability Engineer Tech Lead designs and operates critical infrastructure systems, drives reliability projects across engineering teams, and ensures platform performance and SLAs.
ARE YOU INTERESTED IN JOINING AN INNOVATIVE LOGISTICS TECHNOLOGY COMPANY?
Loadsmart is a growth-stage technology company valued at over $1 billion (a true Tech Unicorn)!
We are a collection of industry veterans and user-centered engineers using innovative technology to fearlessly reinvent the future of freight by helping shippers, brokers, warehouses and carriers to move more with less.
With headquarters in Chicago and a globally distributed remote team, Loadsmart continues to attract top talent committed to driving meaningful change. We seek professionals who embody our core values: curiosity, clarity, results, commitment, and teamwork.
In the SRE, Tech Lead role you will build and maintain the company’s internal platform, driving operational excellence and empowering the entire engineering team. You should have experience in analyzing, proposing, and implementing safer systems and processes. Collaborating closely with engineering squads across platform engineering, you will ensure our applications are both safe and reliable. As a hands-on leader, you will stay actively involved in technical work while collaborating closely with internal stakeholders and engineering squads across the organization to ensure our applications are both safe and reliable.
DEPARTMENT:Â Engineering
LOCATION: Anywhere in Brazil - Remote
WORKING AT LOADSMART:
• Competitive base salaries - we believe in rewarding top talent
• Extremely competitive Equity package - become a shareholder in our company!
• Loadie Time Off - PTO and sick days without a limit
At Loadsmart, we believe our biggest asset is our people. We are proud to be an equal opportunity employer, hiring and developing individuals from diverse backgrounds and experiences to add to our collaborative culture. Loadsmart treats all candidates and employees with respect and does not discriminate in our recruiting, hiring, and promoting processes, including on the basis of race, color, religion, sex, age, sexual orientation, gender identity and/or expression, national origin, veteran status, or disability.
It is the policy of Loadsmart that all offers of employment made shall be contingent upon successful completion of electronic background check(s). These checks will be job-related, consistent with business necessity and conducted by our vendor, pursuant to all applicable laws, rules, policies and procedures of our candidates’ specific locale.
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
Technical Lead oversees GPU infrastructure and Kubernetes-based compute platform development, managing containerization, inference endpoints, and platform observability for Tether Data's AI services.
Join Tether and Shape the Future of Digital Finance
At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.
Innovate with Tether
Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.
But that’s just the beginning:
Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.
Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.
Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.
Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.
Why Join Us?
Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.
If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.
Are you ready to be part of the future?
About the job
Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.
The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.
This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.
Responsibilities
Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.
Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.
Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.
Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.
Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.
Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.
Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.
Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.
Hiring. Complete the platform team and set the technical bar for the engineers who join it.
Must have
Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.
Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.
GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.
High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.
Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.
Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.
HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.
Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.
Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.
A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.
Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.
Excellent written and spoken English. Most partner and leadership work happens in writing.
Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.
Desirable
Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).
Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.
VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).
Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.
Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.
Peer-to-peer or distributed-systems background.
Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.
Important information for candidates
Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:
Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/
Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.
Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.
Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io
We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.
When in doubt, feel free to reach out through our official website.
Technical Lead oversees GPU infrastructure and Kubernetes-based compute platform development, managing distributed systems and platform observability for AI inference services.
Join Tether and Shape the Future of Digital Finance
At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.
Innovate with Tether
Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.
But that’s just the beginning:
Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.
Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.
Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.
Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.
Why Join Us?
Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.
If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.
Are you ready to be part of the future?
About the job
Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.
The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.
This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.
Responsibilities
Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.
Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.
Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.
Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.
Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.
Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.
Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.
Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.
Hiring. Complete the platform team and set the technical bar for the engineers who join it.
Must have
Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.
Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.
GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.
High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.
Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.
Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.
HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.
Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.
Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.
A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.
Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.
Excellent written and spoken English. Most partner and leadership work happens in writing.
Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.
Desirable
Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).
Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.
VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).
Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.
Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.
Peer-to-peer or distributed-systems background.
Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.
Important information for candidates
Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:
Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/
Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.
Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.
Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io
We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.
When in doubt, feel free to reach out through our official website.
Technical Lead manages GPU infrastructure and Kubernetes platform for distributed compute and inference services at Tether Data.
Join Tether and Shape the Future of Digital Finance
At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.
Innovate with Tether
Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.
But that’s just the beginning:
Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.
Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.
Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.
Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.
Why Join Us?
Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.
If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.
Are you ready to be part of the future?
About the job
Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.
The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.
This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.
Responsibilities
Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.
Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.
Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.
Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.
Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.
Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.
Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.
Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.
Hiring. Complete the platform team and set the technical bar for the engineers who join it.
Must have
Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.
Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.
GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.
High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.
Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.
Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.
HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.
Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.
Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.
A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.
Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.
Excellent written and spoken English. Most partner and leadership work happens in writing.
Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.
Desirable
Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).
Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.
VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).
Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.
Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.
Peer-to-peer or distributed-systems background.
Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.
Important information for candidates
Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:
Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/
Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.
Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.
Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io
We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.
When in doubt, feel free to reach out through our official website.
Oversees financial operations, accounting systems, and reporting as the primary financial officer for the organization.
Designs and builds backend infrastructure and systems for an AI-native virtual care platform as a technical lead.
Owns go-to-market strategy, customer acquisition, developer activation, and early-stage business execution to grow an AI gateway platform to 10,000 users.
Headquarters: US
URL: https://openbase.ai/
Head of Growth & Operations — Openbase.ai
Full-time | Remote | Reports directly to the Founder
About Openbase
Openbase.ai is an AI gateway that helps developers and businesses access multiple AI models through one API. We simplify model access, routing, and usage management so teams can focus on building AI products.
Our MVP is ready. We are hiring a hands-on leader to own our go-to-market strategy, day-to-day business execution, and journey from launch to our first 10,000 users.
The opportunity
You will work directly with the founder to turn Openbase into a growing business with repeatable customer acquisition, strong retention, and disciplined operations.
You will own the commercial execution: finding our first customers, testing acquisition channels, improving onboarding, building partnerships, and organizing the people and systems needed to grow.
Our ambition is 10,000 users. Your responsibility is to build a credible, measurable path toward that goal, with clear milestones for activated developers, retained customers, paying accounts, and profitable usage.
What you will own
1. Go-to-market and customer acquisition
2. Developer activation and retention
3. Sales and partnerships
4. Business execution and team management
5. Metrics and commercial performance
What we are looking for
Experience with usage-based pricing, developer relations, technical SEO, and AI model platforms is particularly relevant.
Why join
How to apply
Send your CV or LinkedIn profile, together with brief answers to:
Include an example of a growth experiment or operating system you built. Share only information you are authorized to disclose.
To apply: https://weworkremotely.com/remote-jobs/openbase-llc-head-of-growth-operations