Writing.io Jobs

Find the best remote jobs. Answer a few questions and we'll deploy a powerful assistant to help you search, create alerts, and more.

1 What roles are you open to?

2 Experience level

3 Work style

Did you know? If memory is enabled, Writing.io can remember your job search preferences and help you to improve your resume, craft customized outreach and more.

Support Koast.ai: Head of Technical Support @ Koast.ai

Lead technical support for an AI ad management platform, owning customer onboarding, ticket resolution, QA releases, and engineering triage for media buying teams.

Lead Remote Posted about 6 hours ago We Work Remotely — Programming
What this role involves

Headquarters: Tampa, Florida
URL: https://koast.ai

Head of Technical Support | Full-time | Remote (US hours) | Reports to the CEO

Koast is the AI operating layer for teams that run Meta ads at volume. Agencies, in-house media buying teams, and enterprise buyers managing 10+ ad accounts and launching 50+ ads a week use Koast to launch faster, automate optimizations off real attribution data, and see every account from one Command Center. The Koast Agent is the next layer: an AI that reads the account, spots what matters, and acts inside guardrails.

Our customers are media buyers. When something breaks, it is not a support ticket to them. It is a campaign that did not launch, a budget that did not move, or an account they cannot see. They need someone on the other end who has run ads, knows what they are looking at, and can tell the difference between a Meta API error, a user mistake, and a real bug.

That is this role. You are the anchor between our customers and our engineers. You will onboard every new account, own every ticket, QA every release before it hits production, and make sure engineering hears about patterns, not anecdotes.

What You WIll Own...

Onboarding. Every new customer gets set up right the first time: ad accounts connected, attribution (Hyros, Cometly) wired in, team invited, first launch done together on a call. You will run these sessions yourself, then build the playbook and the in-app flow so the next hundred customers need less of you.

Support, end to end. Intercom is yours. Response times, resolution quality, the help center, the macros, the escalation path. You will answer tickets yourself until the volume justifies a hire, and you will know every customer by name before you delegate a single one.

Triage and the bridge to engineering. You decide what is a bug, what is a feature gap, what is Meta being Meta, and what is user error. Bugs get reproduced, documented with steps and account context, and filed in Shortcut with enough detail that an engineer can fix it without asking you a follow-up question. Patterns get surfaced in the weekly review with numbers: how many customers, how much spend affected, how often.

QA and release quality. Nothing ships to production without you having run it against a real ad account. You will own the QA pass on every release: launch flows, automations, Agent actions, integrations. You know what a broken campaign structure looks like in Ads Manager, and you catch it before a customer does.

System tightness. Meta API connections, token refreshes, webhook health, Whop account provisioning, attribution syncs. You will monitor them (Sentry, PostHog, internal dashboards), notice when something drifts, and get it fixed before it turns into a ticket.

Customer health and retention. You will know which accounts are launching, which have gone quiet, and which are about to churn. You will bring that list to the CEO weekly and you will own the outreach.

What you will do week to week

  • Run onboarding calls for new customers and enterprise accounts. Connect accounts, verify data, launch the first campaign together.
  • Work the Intercom queue. Own first response, resolution, and follow-up.
  • Reproduce every reported bug against a real ad account before it goes to engineering.
  • File and prioritize Shortcut tickets with repro steps, account IDs, screenshots, and expected versus actual behavior.
  • QA each release candidate against a test account and a live customer scenario. Sign off or send it back.
  • Monitor Sentry and PostHog for errors and drop-offs. Flag anything that affects more than one customer the same day.
  • Run the weekly support review: ticket volume by category, time to resolution, top three patterns, customers at risk. One page, numbers first.
  • Write help center articles and Loom walkthroughs for anything you have explained more than twice.

What sets you apart

  • You have run Meta ads. You know Ads Manager, campaign and ad set structure, learning phase, Advantage+, pixel and CAPI setup, and what "the API rejected my creative" actually means. This is a hard requirement, not a preference.
  • You have done support or customer success for a B2B SaaS product and you were the person engineers trusted because your bug reports were correct.
  • You can QA software. You have written test cases, run release checks, and sent a build back because it was not ready.
  • You are calm with a frustrated customer whose launch just failed, and you are direct with an engineer who shipped the thing that failed it.
  • You write tightly. Your tickets, your help articles, and your customer replies are shorter than everyone else's and clearer.
  • You want to build the support function, not just work inside one.

Requirements

  • 3+ years in customer support, customer success, or technical support for a B2B SaaS product, including at least one role where you were the primary bridge to engineering.
  • Hands-on Meta ads experience. You have managed ad accounts, launched campaigns, and troubleshot Ads Manager and API issues. Agency or performance marketing background strongly preferred.
  • Experience with Intercom or a comparable support platform, and a ticketing or project tool (Shortcut, Linear, Jira).
  • Experience running QA or UAT on software releases.
  • Comfortable reading logs, API responses, and error monitoring (Sentry or similar). You do not need to write code. You need to read enough to know where the problem is.
  • Familiarity with attribution tools (Hyros, Cometly, Triple Whale, Northbeam) is a plus.
  • Available full-time during US business hours.

How You Work

Small team, high autonomy. Decisions are made with numbers, written down, and revisited when the numbers change. These are the values we run on. If they sound like you, you will fit right in.

  • Mastery. "Good enough" is where you start, not where you stop. You keep sharpening the craft: tighter repros, faster resolutions, fewer repeat tickets.
  • Push the Boundaries. You take on the ambitious version of the problem. Most support teams answer tickets. You want to build the one that prevents them.
  • Embrace the Challenge. When a customer's launch breaks at 9pm, you are in the thread, not waiting for the morning standup.
  • Responsibility. You own your commitments. If a fix slips, the customer hears it from you first, with a new date attached.
  • Dedication. You dive into the hard, unglamorous problems (an intermittent API error, a broken sync nobody can reproduce) and iterate until they are solved.
  • Collaboration. You write clearly, share what you learn, and make engineering and sales faster for having worked with you.

Why Koast

  • Early-stage impact. You will be one of the first ten full-time employees. The support function, the onboarding playbook, and the QA standard will be the ones you built.
  • AI and ads at scale. You will support some of the highest-volume ad teams on Meta and see what an AI agent does with real budgets every day.
  • Real compensation. Full-time salary, performance bonuses tied to retention and resolution metrics, and paid time off. Bootstrapped and profitable, so the seat is stable and the growth path is real.
  • Remote and flexible. Work from anywhere, on your schedule, during US hours. Just make sure every customer is launching.

How to apply

Record a Loom, five minutes or less, covering these four things in order:

  1. Explain why you're a fit. Screen-share and explain projects youve worked on, ad accounts you have had access to (blur what you need) or companies you helped before. Walk us through the structure and tell us the last thing that broke in it and how you fixed it.
  2. One bug report. Take us through an example of a report you filed that an engineer fixed without asking you a follow-up question. Tell us how you reproduced it.
  3. One angry customer. A time a customer was losing money because of a product issue. What you said in the first reply, what you did next, and what happened.
  4. Where Koast fits. Sixty seconds on what breaks first when a media buyer starts managing 20 accounts from one platform instead of Ads Manager, and how you would catch it before they do.

Email the Loom link and your LinkedIn or resume to jobs@koast.ai with the subject line "Retention is king".

NOTE: DO NOT APPLY IF YOU HAVE NO EXPERIENCE WITH PRODUCTS LIKE THIS / ADVERTISING / CANNOT WORK US HOURS.

No cover letters. We watch the Loom first and read the resume second. Applications without the video are not reviewed.

To apply: https://weworkremotely.com/remote-jobs/koast-ai-head-of-technical-support-koast-ai

Read the full description
Support Mechanical Orchard: Head of Product Support

Leads and scales a product support organization, defining support strategy, building technical support teams, and ensuring customer experience aligns with product excellence.

Lead Remote Posted about 10 hours ago We Work Remotely — Programming
What this role involves

Headquarters: USA (Remote)

Mechanical Orchard is reinventing how the world’s most critical software gets modernized. We’re an applied AI company focused on one of the hardest problems in enterprise technology: rewriting complex legacy systems in a way that is provably correct, low risk, and fast enough to matter. By focusing on system behavior rather than code alone, we turn modernization from a high-stakes, failure-prone effort into a repeatable, confidence-building process that unlocks ongoing innovation.
Joining Mechanical Orchard means working on problems that truly matter—to our customers, their businesses, and the people who rely on these systems every day. We’re building technology and methodology that challenge long-standing industry assumptions, prioritizing proof over promises and progress over theatrics. If you care deeply about quality, rigor, and doing the right thing—even when it’s hard—you’ll find your people here.
About the RoleWe’re hiring a Head of Product Support to build, scale, and lead a world-class Product Support organization.
This is a strategic leadership role responsible for the full support experience across our product offering. You will define our support vision, build the systems and teams that deliver it, and partner closely with Product, Engineering, Delivery, and Sales to ensure customers experience Mechanical Orchard as reliable, responsive, and deeply technical.
You will operate at the intersection of customer experience, product excellence, and operational rigor — shaping how customers interact with our technology and influencing what we build next.
This role may be hired at the Senior Manager or Director level, depending on the candidate’s experience, leadership scope, and demonstrated ability to operate at scale. We are open to calibrating title and level for the right person.
What You’ll Do
Lead and Define the Support Strategy- Own the end-to-end Product Support vision aligned to company and product strategy- Build and mentor a high-performing team of Product Support Specialists, Technical Support Engineers, and Support Operations professionals- Represent Product Support at the leadership level and drive cross-functional initiatives- Establish a culture that is customer-obsessed, data-driven, and technically rigorous
Build Scalable Systems and Operations- Design scalable support workflows, escalation paths, and quality standards- Define and manage SLAs/SLOs aligned with customer and partner expectations- Select and implement support tooling (ticketing, knowledge base, monitoring, automation)- Own support analytics: volume, root causes, trends, friction points, and performance metrics- Partner with Engineering to build automation and internal tooling that reduces repetitive work
Elevate the Customer Experience- Build tight feedback loops between customers and Product- Influence roadmap decisions based on recurring issues and usage patterns- Collaborate with Documentation and Developer Experience teams to improve onboarding and self-serve pathways- Partner with Sales and Delivery on high-priority accounts and escalations- Drive initiatives that proactively reduce support volume through product quality and education
Raise the Technical Bar- Ensure the support organization develops deep product expertise- Build structured training programs to create authoritative technical advisors- Lead post-mortems and embed learnings into product and support systems

Required Requirements

    • 7+ years in Support, Technical Support, Developer Support, or Customer Experience roles, 3+ years leading teams
    • Experience building or scaling support functions in modern tech environments (SaaS, cloud, AI, developer tooling)
    • Strong technical fluency (APIs, debugging, cloud platforms, logs/metrics, dev workflows)
    • Proven ability to operate cross-functionally and influence product direction
    • Experience with modern support tooling and AI-enabled workflows

Preferred Requirements

    • Experience supporting developer-facing or highly technical products
    • Experience designing scalable self-serve support models
    • Background in support operations, quality management, or process engineering
    • Comfort building systems from scratch in ambiguous, high-growth environments
Mechanical Orchard is a remote-first company, with employees working across time zones and locations. This role does not have a requirement for in-office attendance, but is expected to work hours that maximize collaboration with their team (specifics can be discussed during the interview process). There will be travel to company and team meetings, as well as customer visits (1-2 times per quarter, up to ~20% of the time). 
Mechanical Orchard, Inc. is an Equal Opportunity Employer and Prohibits Discrimination and Harassment of Any Kind. Mechanical Orchard, Inc. is committed to the principle of equal employment opportunity for all employees and to providing employees with a work environment free of discrimination and harassment. All employment decisions at Mechanical Orchard, Inc. are based on business needs, job requirements and individual qualifications, without regard to race, color, religion or belief, national, social or ethnic origin, sex (including pregnancy), age, physical, mental or sensory disability, HIV Status, sexual orientation, gender identity and/or expression, marital, civil union or domestic partnership status, past or present military service, family medical history or genetic information, family or parental status, or any other status protected by the laws or regulations in the locations where we operate. Mechanical Orchard, Inc. will not tolerate discrimination or harassment based on any of these characteristics. Mechanical Orchard, Inc. encourages applicants of all ages. Mechanical Orchard, Inc. will provide reasonable accommodation to employees who have protected disabilities consistent with local law.
We look forward to reviewing your application. Thanks!

To apply: https://weworkremotely.com/remote-jobs/mechanical-orchard-head-of-product-support

Read the full description
Project Management Engineering Manager

Manages engineering team, oversees project delivery, and leads technical staff across distributed offices and remote locations.

Lead Remote Posted about 23 hours ago Himalayas
What this role involves
Deputy is a global SaaS workforce management company with hubs in Sydney, Melbourne, San Francisco and London, plus team members working remotely across the United States.
Read the full description
Engineer Site Reliability Engineer, Tech Lead at Loadsmart

Site Reliability Engineer Tech Lead builds and maintains critical infrastructure systems, ensures application reliability and SLAs, and leads reliability projects across engineering squads.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

ARE YOU INTERESTED IN JOINING AN INNOVATIVE LOGISTICS TECHNOLOGY COMPANY?

Loadsmart is a growth-stage technology company valued at over $1 billion (a true Tech Unicorn)!

We are a collection of industry veterans and user-centered engineers using innovative technology to fearlessly reinvent the future of freight by helping shippers, brokers, warehouses and carriers to move more with less.

With headquarters in Chicago and a globally distributed remote team, Loadsmart continues to attract top talent committed to driving meaningful change. We seek professionals who embody our core values: curiosity, clarity, results, commitment, and teamwork.

In the SRE, Tech Lead role you will build and maintain the company’s internal platform, driving operational excellence and empowering the entire engineering team. You should have experience in analyzing, proposing, and implementing safer systems and processes. Collaborating closely with engineering squads across platform engineering, you will ensure our applications are both safe and reliable. As a hands-on leader, you will stay actively involved in technical work while collaborating closely with internal stakeholders and engineering squads across the organization to ensure our applications are both safe and reliable.

DEPARTMENT: Engineering

LOCATION: Anywhere in Brazil - Remote

WHAT YOU GET TO DO

  • Collaborate with and support our creative, tight-knit development team.
  • Design, deploy, and operate Loadsmart’s critical systems while balancing reliability, cost, and agility.
  • Play a key role in driving reliability projects with engineering teams.
  • Utilize your intuitive problem-solving skills and contagious positive attitude to tackle challenging and exciting issues, inspiring those around you.
  • Collect metrics and understand their business impact, encouraging the team to do the same.
  • Perform troubleshooting and root-cause analysis of system operation issues.
  • Be accountable for the platform’s Service Level Agreements and Objectives.
  • Provide infrastructure support during off-hours as needed
  • Take ownership of software infrastructure projects
  • Seek, give, and receive constructive feedback through code and specification reviews.
  • Familiarity with AI agents and agentic workflows, applying AI across the SDLC (AI-assisted coding), LLMs, MCP servers/gateways, and how emerging AI tooling can improve reliability and operations is a plus.

REQUIRED QUALIFICATIONS:

  • 1-3 years leading Reliability Work across multiple engineering squads
  • Over 5 years of experience in Cloud Computing, SRE/DevOps
  • Proven experience collaborating with internal stakeholders across multiple engineering squads
  • Strong project management skills with a demonstrated ability to delegate and mentor team members
  • Proficient in English communication (both written and spoken) to collaborate in an international team with native and non-native English speakers
  • Detail-oriented with high initiative and self-motivation
  • Strong understanding of software engineering principles and how systems work under the hood
  • In-depth knowledge of modern networking and operating systems
  • Proficiency in AWS, cloud environments, containers, Kubernetes, Docker, and DevOps engineering, including managing tests and CI/CD pipelines
  • Familiarity with automation tools and provisioners like Terraform, Ansible, or Chef
  • Solid troubleshooting and system engineering experience in UNIX/Linux production environments
  • Experience with monitoring, alerting, and incident management
  • Proficiency in automating tasks with scripting languages like Python, Bash, etc
  • Experience or exposure to PostgreSQL and DBA responsibilities is a plus
  • Fluent in English (both written and spoken); comfortable interacting with native English speakers daily.

WORKING AT LOADSMART:

• Competitive base salaries - we believe in rewarding top talent

• Extremely competitive Equity package - become a shareholder in our company!

• Loadie Time Off - PTO and sick days without a limit

At Loadsmart, we believe our biggest asset is our people. We are proud to be an equal opportunity employer, hiring and developing individuals from diverse backgrounds and experiences to add to our collaborative culture. Loadsmart treats all candidates and employees with respect and does not discriminate in our recruiting, hiring, and promoting processes, including on the basis of race, color, religion, sex, age, sexual orientation, gender identity and/or expression, national origin, veteran status, or disability.

It is the policy of Loadsmart that all offers of employment made shall be contingent upon successful completion of electronic background check(s). These checks will be job-related, consistent with business necessity and conducted by our vendor, pursuant to all applicable laws, rules, policies and procedures of our candidates’ specific locale.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Read the full description
Engineer Technical Lead - GPU Infrastructure at Tether.io

Lead GPU infrastructure engineering for a Kubernetes-based compute platform, managing containerized GPU resources and inference endpoints at scale.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Description

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.

The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.

This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.

Responsibilities

Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.

Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.

Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.

Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.

Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.

Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.

Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.

Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.

Hiring. Complete the platform team and set the technical bar for the engineers who join it.

Requirements

Must have

  • Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

  • Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.

  • GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.

  • High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.

  • Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.

  • Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.

  • HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.

  • Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.

  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.

  • A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.

  • Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.

  • Excellent written and spoken English. Most partner and leadership work happens in writing.

  • Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.

Desirable

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).

  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.

  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).

  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.

  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.

  • Peer-to-peer or distributed-systems background.

  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Important information for candidates

Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Read the full description
Engineer Technical Lead - GPU Infrastructure at Tether.io

Technical Lead oversees GPU infrastructure and Kubernetes platform development for Tether Data's managed inference and compute services.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Description

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.

The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.

This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.

Responsibilities

Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.

Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.

Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.

Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.

Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.

Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.

Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.

Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.

Hiring. Complete the platform team and set the technical bar for the engineers who join it.

Requirements

Must have

  • Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

  • Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.

  • GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.

  • High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.

  • Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.

  • Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.

  • HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.

  • Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.

  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.

  • A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.

  • Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.

  • Excellent written and spoken English. Most partner and leadership work happens in writing.

  • Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.

Desirable

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).

  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.

  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).

  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.

  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.

  • Peer-to-peer or distributed-systems background.

  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Important information for candidates

Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Read the full description
Engineer Technical Lead - GPU Infrastructure at Tether.io

Technical Lead manages GPU infrastructure and Kubernetes-based compute platform, overseeing architecture, performance, and team delivery for a managed inference service.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Description

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.

The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.

This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.

Responsibilities

Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.

Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.

Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.

Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.

Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.

Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.

Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.

Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.

Hiring. Complete the platform team and set the technical bar for the engineers who join it.

Requirements

Must have

  • Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

  • Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.

  • GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.

  • High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.

  • Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.

  • Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.

  • HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.

  • Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.

  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.

  • A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.

  • Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.

  • Excellent written and spoken English. Most partner and leadership work happens in writing.

  • Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.

Desirable

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).

  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.

  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).

  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.

  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.

  • Peer-to-peer or distributed-systems background.

  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Important information for candidates

Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Read the full description
Sales Sales Manager, Revenue Generation - EMEA at BlueOptima

Leads a team of Account Executives selling enterprise SaaS software productivity solutions to CIOs, CTOs, and engineering leaders across EMEA, managing complex deals and coaching team performance.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Company Description

BlueOptima gives CIOs, CTOs and CFOs objective, code-level visibility into software productivity, AI ROI, and engineering cost. This is a problem that until recently could not be measured at all. With AI-generated code now flowing into enterprise codebases at scale, that measurement gap has become a board-level priority.

10B+ code revisions analysed · 800,000+ developers on platform · 9 of the top 16 universal banks

The business is profitable, bootstrapped, and 20 years deep in enterprise relationships. You are joining a sales organisation that is mid-build. The foundations are in place, the product has genuine enterprise traction, and there is real room to grow fast.

Buyers are CIOs, CTOs, VP Engineering and CFO stakeholders at large software engineering organisations. The conversation is about developer productivity, engineering cost, and AI ROI. These are budget-backed problems that sit at the top of the technology agenda right now. You are not convincing anyone the problem exists.

Job Description

You will lead a team of Account Executives working on enterprise accounts across the UK and EMEA, with deal sizes typically in the ÂŁ100k to ÂŁ200k ARR range. You will be cultivating a culture of success and continuous improvement through your coaching. Your leadership will guide the team through complex sales cycles and ambitious growth targets.

You will use sales data and team performance metrics to identify opportunities for improvement. Implement actionable insights to enhance sales tactics and team productivity, leveraging coaching frameworks to work side by side with your team members.

You will work in close partnership with Marketing and SDR leadership to identify new pipeline generation opportunities, and with Customer Success leadership to surface expansion opportunities within the install base.

Qualifications

Required

  • Enterprise SaaS sales leadership: At least 3 years managing sales teams closing complex enterprise deals in the ÂŁ100k to ÂŁ200k+ ARR range.
  • Engineering domain experience: Direct experience selling into software engineering organisations is required. This means you have sold developer tooling, engineering analytics, observability, SDLC platforms, or similar products and can hold a credible technical conversation with VP Engineering and CTO buyers without leaning on a solutions engineer. You do not need to have been a developer, but you do need to understand how software gets built.
  • Executive-level selling: Comfortable engaging CIO, CTO, VP Engineering and CFO buyers. You can hold the room with credibility and navigate multi-stakeholder deal structures.
  • Hands-on coaching style: Demonstrated ability to develop your team through direct involvement in sales activities, including joint calls and co-selling, to ensure team members reach their potential.
  • Data-driven decision making: Highly analytical, you leverage pipeline and performance data to make informed decisions and adapt quickly to capitalise on new opportunities.
  • Structured sales methodology: Exposure to value-based selling methodologies and the MEDDPICC discovery framework in multi-stakeholder environments.

Valued

  • Experience with AI products, AI governance, or engineering productivity measurement.
  • Background in land-and-expand motions within large enterprise accounts.
  • Open to using AI to amplify your own skills and strengthen your work, demonstrating curiosity, willingness to learn, and sound judgement in applying AI responsibly to improve efficiency and impact.

Additional Information

Why This Role Makes Sense Right Now

AI governance and engineering ROI measurement have moved from nice-to-have to board priority in the last 18 months. The buyers are funded, the problem is urgent, and BlueOptima is one of the very few companies with a production-grade answer. For a sales leader who wants to work with a team closing real enterprise deals, not chasing SMB volume, the timing is good.

Where This Leads

The natural progression from this role is into second-line leadership. That path is defined and is based on performance and the growth of the company, not tenure.

Benefits we offer

  • London HQ, 3 days in office, 2 remote
  • 32 days holiday including bank holidays
  • 12 weeks paid maternity and paternity leave
  • 4 weeks per year flexible remote working from anywhere
  • Annual leave purchase up to 10 extra days
  • Work from home equipment allowance
  • Pet friendly office
  • Sponsored learning opportunities
  • Cycle to work scheme
  • Team socials

Stay connected with us on LinkedIn or keep an eye on our career page for future opportunities!

Read the full description
Operations Staff Site Reliability Engineer at Fingerprint

Define and implement SLIs/SLOs, establish error budgets, strengthen incident response processes, and coach engineering teams on reliability best practices across the platform.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Fingerprint empowers enterprises to detect and stop online fraud with the world’s most accurate device intelligence.  We lead our industry with bleeding-edge identification capabilities and work on turning new ideas and discoveries in the fraud detection space into reality. Our customers range from innovative startups to leading enterprise companies, including Plaid, Dropbox, and Booking.com.

Fingerprint is a globally dispersed, 100% remote company. We were named on on the 2026 Forbes Best Startup Employers list and ranked #803 on the 2026 Inc. 5000 list of America’s fastest-growing private companies.

We have raised $77M and are backed by Craft Ventures (Tesla, Facebook, Airbnb ), Nexus Venture Partners ( Postman, Apollo.io, MinIO, Druva) and Uncorrelated Ventures ( Redis, Rollbar,  Gradle).

About the role

You will be Fingerprint’s first dedicated Site Reliability Engineer. You will work alongside our Architect on the shape of the platform, with Cloud Platform on the infrastructure that runs it, and with every product team on how they operate what they own — without direct reports. The mandate has three parts. First, make reliability measurable: define SLIs and SLOs for our critical paths, get teams to own them, and make error budgets the shared language for prioritizing reliability against features. Second, raise the operational bar: strengthen incident response, postmortem quality, alerting, and change safety so that we find out first and repeat incidents stop repeating. Third, build the mindset: coach teams to design for failure, test for it deliberately, and treat operability as part of done — so the practices outlive your involvement in any single team.

You report directly to the VP of Engineering. That placement is deliberate: reliability standards need to apply evenly across eight engineering groups, and you need the neutrality to hold every team, including infrastructure, to the same bar.

What you’ll do

Make reliability measurable

  • Define SLIs and SLOs for Fingerprint’s critical request paths (identification, events, server APIs, client agents) with the teams that own them; make them visible, reviewed, and tied to decisions.
  • Introduce error budgets as the mechanism for balancing reliability investment against feature work, and coach EMs and Staff engineers on using them.
  • Own the reliability metrics that leadership uses to judge progress; be the No Nonsense voice on whether we are actually getting better.

Raise the operational bar

  • Strengthen the incident lifecycle end to end: detection, response, communication, postmortem quality, and follow-up completion. Make the postmortem the most useful document a team writes.
  • Close the “customers find out before we do” gap: drive alert quality, correctness anomaly detection, and escalation design across teams, working with Cloud Platform on shared tooling.
  • Lead reliability reviews for high-risk changes and new services (production readiness, capacity, failure modes, rollback), teaching through review rather than gatekeeping.
  • Introduce deliberate failure testing (game days, chaos exercises) in staging first, then production, to discover gaps and safe limits before customers do.

Build the SRE mindset in teams

  • Embed with teams for time-boxed engagements: pair on their hardest reliability problems, leave behind better practices and a stronger owner, then move on.
  • Develop Staff and Lead engineers as reliability leaders in their own groups — the goal is that every team has someone who thinks like an SRE.
  • Codify practices that stick: production readiness checklists, on-call standards, runbook quality, change safety norms. Make them lightweight enough that teams choose to use them.
  • Partner with the Architect and tech leads so reliability is designed in, not retrofitted.

Stay hands-on

  • Dig into production during incidents and investigations. Write tooling, dashboards, and reference implementations. Be credible with the engineers you are asking to change how they work.
  • Lead AI adoption in reliability practice: set norms for AI-assisted incident investigation, postmortem analysis, runbook authoring, and observability tooling across teams, and shape our runbooks, alerts, and operational data so AI agents can safely help diagnose and operate our systems alongside engineers.

What we’re looking for

  • 10+ years of engineering experience, with 3+ years as an SRE, production engineer, or reliability-focused Staff engineer operating across multiple teams — you have owned reliability for a platform, not just for a service you built.
  • Deep experience with SLI/SLO design and error budgets in practice, including the hard part: getting product teams to adopt and act on them.
  • Strong incident leadership: you have run incident response and postmortems for high-severity, customer-facing incidents and materially improved how an organization learns from them.
  • Hands-on depth in distributed systems failure modes — cache/database saturation and cascading failure, retry storms, capacity limits, degradation and load shedding — in a high-throughput, low-latency environment. Fluent in Kubernetes, AWS, and modern observability tooling (Datadog or equivalent).
  • Comfortable reading and writing production code (Go, TypeScript, or similar) and infrastructure as code. You can ship a fix, not just recommend one.
  • Track record of leading through influence: you have changed how teams you did not manage operate, and can explain how adoption actually happened.
  • Teacher’s instinct. You have coached engineers into owning reliability and can point to practices that persisted after you stepped back.
  • Exceptional written communication. You make incidents, risks, and trade-offs legible to engineers and executives alike, and you default to async, documented decision-making.
  • AI-native by default. You use AI tools as a normal part of how you investigate incidents, analyze telemetry, write runbooks and postmortems, and build tooling — and you have opinions, from experience, about where they accelerate reliability work and where they don’t yet.
  • Operate for an AI-assisted org. You think about how runbooks, alerts, dashboards, and operational data should be structured so that both humans and AI agents can diagnose and act on them safely — legible signals, clear ownership, strong guardrails on automated change.
  • Pragmatism over purity. You know that reliability competes with delivery, and you can make the case for the right investment at the right time — and say when a risk is acceptable.

Nice to have

  • Experience in fraud detection, identity, payments, or other adversarial, real-time domains.
  • Multi-region, cell-based, or failure-isolation architecture experience.
  • Experience with Elasticsearch, Redis, DynamoDB, or Kafka at scale, including their failure modes.
  • Familiarity with FinOps and the reliability/cost trade-off in cloud infrastructure.

Compensation & Transparency

For US-based employees, the cash compensation range for this role is $177,000 – $240,000. We set standard ranges for all US roles based on function, level, and geographic location, benchmarked against similar stage growth companies. To comply with local legislation and provide greater transparency, we share salary ranges on all job postings. However, these ranges are specific to the hiring location and may differ within or outside the US. Offers vary depending on, but not limited to, relevant experience, education, certifications/licenses, skills, training, and market conditions.

Due to regulatory and security reasons, there’s a small number of countries where we cannot have Fingerprint teammates based. Additionally, because Fingerprint is an all-remote company and people can join our workforce from almost any country, we do not sponsor visas. Fingerprint teammates need to be authorized to work from their home location.

We are dedicated to creating an inclusive work environment for everyone. We embrace and celebrate the unique experiences, perspectives and cultural backgrounds that each employee brings to our workplace. Fingerprint strives to foster an environment where our employees feel respected, valued and empowered, and our team members are at the forefront in helping us promote and sustain an inclusive workplace. We highly encourage people from underrepresented groups in tech to apply.

If you are applying as a resident of California, please read our CCPA notice here.

If you are applying as a resident of the EU, please read our GDPR notice here.

  • We have noticed a rise in recruiting impersonations across the industry, where scammers attempt to access candidates’ personal and financial information through fake interviews and offers. All Fingerprint recruiting email communications will always come from the @fingerprint.com domain. Any outreach claiming to be from Fingerprint via other sources should be ignored.*

Due to regulatory and security reasons, there’s a small number of countries where we cannot have Fingerprint teammates based. Additionally, because Fingerprint is an all-remote company and people can join our workforce from almost any country, we do not sponsor visas. Fingerprint teammates need to be authorized to work from their home location.

We are dedicated to creating an inclusive work environment for everyone. We embrace and celebrate the unique experiences, perspectives and cultural backgrounds that each employee brings to our workplace. Fingerprint strives to foster an environment where our employees feel respected, valued and empowered, and our team members are at the forefront in helping us promote and sustain an inclusive workplace. We highly encourage people from underrepresented groups in tech to apply.

If you are applying as a resident of California, please read our CCPA notice here.

If you are applying as a resident of the EU, please read our GDPR notice here.

**We have noticed a rise in recruiting impersonations across the industry, where scammers attempt to access candidates’ personal and financial information through fake interviews and offers. All Fingerprint recruiting email communications will always come from the @fingerprint.com domain. Any outreach claiming to be from Fingerprint via other sources should be ignored.

Read the full description
Engineer Software Architect at Fingerprint

Lead cross-cutting architecture decisions across engineering teams, designing scalable systems and setting technical standards for a fraud detection platform.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Fingerprint empowers enterprises to detect and stop online fraud with the world’s most accurate device intelligence.  We lead our industry with bleeding-edge identification capabilities and work on turning new ideas and discoveries in the fraud detection space into reality. Our customers range from innovative startups to leading enterprise companies, including Plaid, Dropbox, and Booking.com.

Fingerprint is a globally dispersed, 100% remote company. We were named on on the 2026 Forbes Best Startup Employers list and ranked #803 on the 2026 Inc. 5000 list of America’s fastest-growing private companies.

We have raised $77M and are backed by Craft Ventures (Tesla, Facebook, Airbnb ), Nexus Venture Partners ( Postman, Apollo.io, MinIO, Druva) and Uncorrelated Ventures ( Redis, Rollbar,  Gradle).

About the role

You will be Fingerprint’s first dedicated Architect. You will lead cross-cutting architecture — how our systems fit together and where they need to go — working alongside the Staff and Lead engineers who already shape them, without direct reports. The mandate has three parts. First, go deep: build the end-to-end picture nobody currently has time to hold, identify where the platform will strain as we scale, and set the direction to address it. Second, raise the bar: strengthen and extend our design review, API, service, and reliability practices so eight engineering groups can move fast without stepping on each other. Third, make it durable: give cross-cutting architecture work consistent cadence, a durable decision record, and follow-through, so strong individual judgment compounds into platform-level outcomes.

You report directly to the VP of Engineering. That placement is deliberate — you need neutrality across groups so the standards you set with teams are adopted everywhere.

What you’ll do

Own the end-to-end architecture

  • Build and maintain the end-to-end architecture picture — data flows, service boundaries, ownership seams — and keep it current as the org changes.
  • Identify systemic scaling, reliability, and cost risks across the pipeline (ingestion, identification, signals, delivery) and drive architectural changes to address them before they become incidents.
  • Set direction for API design and service patterns across REST, SDK, and MCP surfaces so customers see one coherent platform.
  • Shape the roadmap for foundational initiatives (e.g., cell architecture, multi-region readiness, rate limiting, feature flag infrastructure) and lead the hardest designs yourself.

Lead cross-cutting architecture

  • Serve as full-time technical lead across the engineering org: set the agenda, drive decisions to closure, own the decision record, and follow through on adoption across teams.
  • Partner with Engineering leadership and the Staff/Lead engineers across teams to prioritize cross-cutting work and connect it to team roadmaps.
  • Mentor Staff-level engineers on system design, technical writing, and influence — raise the bar for what “Staff” means at Fingerprint.

Strengthen and extend best practices

  • Build on the design review practices we already have (RFCs, cross-team design reviews) and make them more consistent, lighter-weight, and more useful, so teams choose to use them.
  • Codify and extend standards for observability, quality, security-by-design, and API evolution/versioning, working with Cloud Platform, Security, and product teams.
  • Review and sign off on critical designs; teach through review rather than gatekeeping.

Set us up for scale

  • Translate business trajectory (Tier 1 enterprise customers, new product launches, AI/MCP workflows) into a multi-quarter technical strategy and sequence the work.
  • Advise Engineering leadership and executive team on build/buy, platform investment, and technical risk in plain language.
  • Lead AI adoption in engineering practice: set norms for AI-assisted design, coding, and review across teams, and shape our architecture and documentation so AI agents can work in our systems as effectively as engineers do.
  • Stay hands-on: prototype, write reference implementations and documentation, and dig into production when the problem demands it.

What we’re looking for

  • 10+ years of software engineering experience, including 3+ years operating as an Architect for a 100+ Engineering org — you have owned architecture for a platform, not just a service.
  • Track record of leading through influence: you have led senior engineers you didn’t manage, run design review or architecture forums that people actually used, and driven adoption of standards across an org.
  • Deep expertise in distributed systems and high-throughput, low-latency backend architecture (Go or similar systems language; Kubernetes/AWS; event-driven and data-intensive systems). Fluency across adjacent layers — client SDKs, data pipelines, ML serving — is a strong plus.
  • Experience designing public APIs and SDKs at scale, including versioning, backward compatibility, and multi-surface consistency.
  • Demonstrated ability to anticipate systemic risk and act on it — you can point to failures you prevented, not just ones you fixed.
  • Exceptional written communication. You make complex decisions legible to engineers and executives alike, and you default to async, documented decision-making.
  • AI-native by default. You use AI coding and reasoning tools as a normal part of how you design, prototype, review, and write — and you have opinions, from experience, about where they accelerate engineering work and where they don’t yet.
  • Architect for an AI-assisted org. You think about how codebases, documentation, service boundaries, and APIs should be shaped so that both humans and AI agents can work in them safely — legible structure, strong contracts, automated verification.
  • Pragmatism over purity. You balance long-term architecture with delivery pressure and know when “good enough” is the right call.
  • Comfortable being the first: you have stepped into a dedicated role where the work was previously shared across a group, earned the trust of the people already doing it, and made them more effective rather than displacing them.

Nice to have

  • Experience in fraud detection, identity, device intelligence, or other adversarial domains.
  • Multi-region / cell-based architecture experience.
  • Experience with MCP, LLM tool integrations, or agent-facing API design.
  • Familiarity with ClickHouse, Kafka, or similar high-volume data infrastructure.

Compensation & Transparency

For US-based employees, the cash compensation range for this role is $177,000 – $240,000. We set standard ranges for all US roles based on function, level, and geographic location, benchmarked against similar stage growth companies. To comply with local legislation and provide greater transparency, we share salary ranges on all job postings. However, these ranges are specific to the hiring location and may differ within or outside the US. Offers vary depending on, but not limited to, relevant experience, education, certifications/licenses, skills, training, and market conditions.

Due to regulatory and security reasons, there’s a small number of countries where we cannot have Fingerprint teammates based. Additionally, because Fingerprint is an all-remote company and people can join our workforce from almost any country, we do not sponsor visas. Fingerprint teammates need to be authorized to work from their home location.

We are dedicated to creating an inclusive work environment for everyone. We embrace and celebrate the unique experiences, perspectives and cultural backgrounds that each employee brings to our workplace. Fingerprint strives to foster an environment where our employees feel respected, valued and empowered, and our team members are at the forefront in helping us promote and sustain an inclusive workplace. We highly encourage people from underrepresented groups in tech to apply.

If you are applying as a resident of California, please read our CCPA notice here.

If you are applying as a resident of the EU, please read our GDPR notice here.

  • We have noticed a rise in recruiting impersonations across the industry, where scammers attempt to access candidates’ personal and financial information through fake interviews and offers. All Fingerprint recruiting email communications will always come from the @fingerprint.com domain. Any outreach claiming to be from Fingerprint via other sources should be ignored.*

Due to regulatory and security reasons, there’s a small number of countries where we cannot have Fingerprint teammates based. Additionally, because Fingerprint is an all-remote company and people can join our workforce from almost any country, we do not sponsor visas. Fingerprint teammates need to be authorized to work from their home location.

We are dedicated to creating an inclusive work environment for everyone. We embrace and celebrate the unique experiences, perspectives and cultural backgrounds that each employee brings to our workplace. Fingerprint strives to foster an environment where our employees feel respected, valued and empowered, and our team members are at the forefront in helping us promote and sustain an inclusive workplace. We highly encourage people from underrepresented groups in tech to apply.

If you are applying as a resident of California, please read our CCPA notice here.

If you are applying as a resident of the EU, please read our GDPR notice here.

**We have noticed a rise in recruiting impersonations across the industry, where scammers attempt to access candidates’ personal and financial information through fake interviews and offers. All Fingerprint recruiting email communications will always come from the @fingerprint.com domain. Any outreach claiming to be from Fingerprint via other sources should be ignored.

Read the full description
Engineer Technical Lead - GPU Infrastructure at Tether.io

Technical Lead oversees GPU infrastructure and Kubernetes-based compute platform development, managing architecture and team for Tether Data's managed inference services.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Description

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.

The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.

This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.

Responsibilities

Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.

Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.

Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.

Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.

Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.

Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.

Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.

Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.

Hiring. Complete the platform team and set the technical bar for the engineers who join it.

Requirements

Must have

  • Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

  • Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.

  • GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.

  • High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.

  • Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.

  • Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.

  • HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.

  • Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.

  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.

  • A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.

  • Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.

  • Excellent written and spoken English. Most partner and leadership work happens in writing.

  • Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.

Desirable

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).

  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.

  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).

  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.

  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.

  • Peer-to-peer or distributed-systems background.

  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Important information for candidates

Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Read the full description
Engineer Technical Lead - GPU Infrastructure at Tether.io

Technical Lead designs and manages GPU infrastructure and Kubernetes-based compute platform for AI inference and managed services.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Description

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.

The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.

This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.

Responsibilities

Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.

Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.

Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.

Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.

Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.

Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.

Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.

Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.

Hiring. Complete the platform team and set the technical bar for the engineers who join it.

Requirements

Must have

  • Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

  • Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.

  • GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.

  • High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.

  • Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.

  • Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.

  • HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.

  • Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.

  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.

  • A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.

  • Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.

  • Excellent written and spoken English. Most partner and leadership work happens in writing.

  • Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.

Desirable

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).

  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.

  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).

  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.

  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.

  • Peer-to-peer or distributed-systems background.

  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Important information for candidates

Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Read the full description
Engineer Technical Lead - GPU Infrastructure at Tether.io

Technical Lead oversees GPU infrastructure and Kubernetes-based platform development, managing compute resources and managed inference endpoints for Tether Data's AI infrastructure.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Description

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.

The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.

This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.

Responsibilities

Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.

Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.

Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.

Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.

Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.

Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.

Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.

Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.

Hiring. Complete the platform team and set the technical bar for the engineers who join it.

Requirements

Must have

  • Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

  • Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.

  • GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.

  • High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.

  • Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.

  • Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.

  • HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.

  • Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.

  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.

  • A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.

  • Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.

  • Excellent written and spoken English. Most partner and leadership work happens in writing.

  • Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.

Desirable

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).

  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.

  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).

  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.

  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.

  • Peer-to-peer or distributed-systems background.

  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Important information for candidates

Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Read the full description
Engineer Site Reliability Engineer, Tech Lead at Loadsmart

Site Reliability Engineer Tech Lead designs and operates critical infrastructure systems, drives reliability projects across engineering teams, and ensures platform performance and SLAs.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

ARE YOU INTERESTED IN JOINING AN INNOVATIVE LOGISTICS TECHNOLOGY COMPANY?

Loadsmart is a growth-stage technology company valued at over $1 billion (a true Tech Unicorn)!

We are a collection of industry veterans and user-centered engineers using innovative technology to fearlessly reinvent the future of freight by helping shippers, brokers, warehouses and carriers to move more with less.

With headquarters in Chicago and a globally distributed remote team, Loadsmart continues to attract top talent committed to driving meaningful change. We seek professionals who embody our core values: curiosity, clarity, results, commitment, and teamwork.

In the SRE, Tech Lead role you will build and maintain the company’s internal platform, driving operational excellence and empowering the entire engineering team. You should have experience in analyzing, proposing, and implementing safer systems and processes. Collaborating closely with engineering squads across platform engineering, you will ensure our applications are both safe and reliable. As a hands-on leader, you will stay actively involved in technical work while collaborating closely with internal stakeholders and engineering squads across the organization to ensure our applications are both safe and reliable.

DEPARTMENT: Engineering

LOCATION: Anywhere in Brazil - Remote

WHAT YOU GET TO DO

  • Collaborate with and support our creative, tight-knit development team.
  • Design, deploy, and operate Loadsmart’s critical systems while balancing reliability, cost, and agility.
  • Play a key role in driving reliability projects with engineering teams.
  • Utilize your intuitive problem-solving skills and contagious positive attitude to tackle challenging and exciting issues, inspiring those around you.
  • Collect metrics and understand their business impact, encouraging the team to do the same.
  • Perform troubleshooting and root-cause analysis of system operation issues.
  • Be accountable for the platform’s Service Level Agreements and Objectives.
  • Provide infrastructure support during off-hours as needed
  • Take ownership of software infrastructure projects
  • Seek, give, and receive constructive feedback through code and specification reviews.
  • Familiarity with AI agents and agentic workflows, applying AI across the SDLC (AI-assisted coding), LLMs, MCP servers/gateways, and how emerging AI tooling can improve reliability and operations is a plus.

REQUIRED QUALIFICATIONS:

  • 1-3 years leading Reliability Work across multiple engineering squads
  • Over 5 years of experience in Cloud Computing, SRE/DevOps
  • Proven experience collaborating with internal stakeholders across multiple engineering squads
  • Strong project management skills with a demonstrated ability to delegate and mentor team members
  • Proficient in English communication (both written and spoken) to collaborate in an international team with native and non-native English speakers
  • Detail-oriented with high initiative and self-motivation
  • Strong understanding of software engineering principles and how systems work under the hood
  • In-depth knowledge of modern networking and operating systems
  • Proficiency in AWS, cloud environments, containers, Kubernetes, Docker, and DevOps engineering, including managing tests and CI/CD pipelines
  • Familiarity with automation tools and provisioners like Terraform, Ansible, or Chef
  • Solid troubleshooting and system engineering experience in UNIX/Linux production environments
  • Experience with monitoring, alerting, and incident management
  • Proficiency in automating tasks with scripting languages like Python, Bash, etc
  • Experience or exposure to PostgreSQL and DBA responsibilities is a plus
  • Fluent in English (both written and spoken); comfortable interacting with native English speakers daily.

WORKING AT LOADSMART:

• Competitive base salaries - we believe in rewarding top talent

• Extremely competitive Equity package - become a shareholder in our company!

• Loadie Time Off - PTO and sick days without a limit

At Loadsmart, we believe our biggest asset is our people. We are proud to be an equal opportunity employer, hiring and developing individuals from diverse backgrounds and experiences to add to our collaborative culture. Loadsmart treats all candidates and employees with respect and does not discriminate in our recruiting, hiring, and promoting processes, including on the basis of race, color, religion, sex, age, sexual orientation, gender identity and/or expression, national origin, veteran status, or disability.

It is the policy of Loadsmart that all offers of employment made shall be contingent upon successful completion of electronic background check(s). These checks will be job-related, consistent with business necessity and conducted by our vendor, pursuant to all applicable laws, rules, policies and procedures of our candidates’ specific locale.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Read the full description
Engineer Technical Lead - GPU Infrastructure at Tether.io

Technical Lead oversees GPU infrastructure and Kubernetes-based compute platform development, managing containerization, inference endpoints, and platform observability for Tether Data's AI services.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Description

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.

The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.

This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.

Responsibilities

Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.

Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.

Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.

Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.

Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.

Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.

Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.

Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.

Hiring. Complete the platform team and set the technical bar for the engineers who join it.

Requirements

Must have

  • Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

  • Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.

  • GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.

  • High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.

  • Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.

  • Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.

  • HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.

  • Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.

  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.

  • A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.

  • Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.

  • Excellent written and spoken English. Most partner and leadership work happens in writing.

  • Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.

Desirable

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).

  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.

  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).

  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.

  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.

  • Peer-to-peer or distributed-systems background.

  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Important information for candidates

Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Read the full description
Engineer Technical Lead - GPU Infrastructure at Tether.io

Technical Lead oversees GPU infrastructure and Kubernetes-based compute platform development, managing distributed systems and platform observability for AI inference services.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Description

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.

The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.

This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.

Responsibilities

Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.

Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.

Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.

Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.

Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.

Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.

Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.

Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.

Hiring. Complete the platform team and set the technical bar for the engineers who join it.

Requirements

Must have

  • Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

  • Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.

  • GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.

  • High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.

  • Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.

  • Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.

  • HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.

  • Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.

  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.

  • A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.

  • Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.

  • Excellent written and spoken English. Most partner and leadership work happens in writing.

  • Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.

Desirable

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).

  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.

  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).

  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.

  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.

  • Peer-to-peer or distributed-systems background.

  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Important information for candidates

Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Read the full description
Engineer Technical Lead - GPU Infrastructure at Tether.io

Technical Lead manages GPU infrastructure and Kubernetes platform for distributed compute and inference services at Tether Data.

Lead Remote Posted about 24 hours ago RemoteFirstJobs Product
What this role involves

Description

Join Tether and Shape the Future of Digital Finance

At Tether, we’re not just building products, we’re pioneering a global financial revolution. Our cutting-edge solutions empower businesses—from exchanges and wallets to payment processors and ATMs—to seamlessly integrate reserve-backed tokens across blockchains. By harnessing the power of blockchain technology, Tether enables you to store, send, and receive digital tokens instantly, securely, and globally, all at a fraction of the cost. Transparency is the bedrock of everything we do, ensuring trust in every transaction.

Innovate with Tether

Tether Finance: Our innovative product suite features the world’s most trusted stablecoin, USDT, relied upon by hundreds of millions worldwide, alongside pioneering digital asset tokenization services.

But that’s just the beginning:

Tether Power: Driving sustainable growth, our energy solutions optimize excess power for Bitcoin mining using eco-friendly practices in state-of-the-art, geo-diverse facilities.

Tether Data: Fueling breakthroughs in AI and peer-to-peer technology, we reduce infrastructure costs and enhance global communications with cutting-edge solutions like KEET, our flagship app that redefines secure and private data sharing.

Tether Education: Democratizing access to top-tier digital learning, we empower individuals to thrive in the digital and gig economies, driving global growth and opportunity.

Tether Evolution: At the intersection of technology and human potential, we are pushing the boundaries of what is possible, crafting a future where innovation and human capabilities merge in powerful, unprecedented ways.

Why Join Us?

Our team is a global talent powerhouse, working remotely from every corner of the world. If you’re passionate about making a mark in the fintech space, this is your opportunity to collaborate with some of the brightest minds, pushing boundaries and setting new standards. We’ve grown fast, stayed lean, and secured our place as a leader in the industry.

If you have excellent English communication skills and are ready to contribute to the most innovative platform on the planet, Tether is the place for you.

Are you ready to be part of the future?

About the job

Cosmic AC is Tether Data’s GPU compute and managed inference platform: GPU containers, managed inference endpoints and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that.

The Technical Lead owns the architecture and delivery of that stack and leads the engineering team building it: about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers, and is the primary technical interface to our infrastructure partners.

This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not a management-only role.

Responsibilities

Architecture. Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.

Team leadership. Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation: engineering standards, code and design review, release gates, one-to-ones, growth and performance input.

Bare-metal GPU scheduling layer. Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.

Kubernetes control plane and GPU enablement. Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations: upgrades, backup and recovery, node replacement.

Managed inference at scale. Serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability; confidential-compute-capable capacity for sensitive workloads.

Observability and operations. Metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers; incident response and post-incident review; an on-call model a small team can sustain.

Partners and vendors. Primary technical interface to infrastructure partners and vendors: turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.

Internal consumers. Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.

Hiring. Complete the platform team and set the technical bar for the engineers who join it.

Requirements

Must have

  • Experience. Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.

  • Slurm at scale, hands on. Has run slurmctld and slurmdbd for real users: partitions, QoS and priority, accounting, prolog and epilog, node health scripting, upgrades with jobs on the system. Ideally has operated an HPC or GPU training cluster for a research population.

  • GPU fleet operation on bare metal. NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance.

  • High-performance interconnects. InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.

  • Linux systems depth. Kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, performance tuning for compute-heavy workloads.

  • Production Kubernetes operation, not just deployment: control plane, upgrades, CNI and CSI, operators and custom controllers, multi-tenancy design.

  • HPC storage and data movement. Shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, distributing large model weights and datasets across many nodes.

  • Observability and operations. Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.

  • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them. Not a feature-development requirement.

  • A shipped platform with real users. A multi-tenant IaaS or PaaS, or a research computing service: resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.

  • Leadership that stays in the code. People management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.

  • Excellent written and spoken English. Most partner and leadership work happens in writing.

  • Location. Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.

Desirable

  • Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).

  • Modern serving stacks (vLLM, SGLang, TensorRT-LLM): parallelism strategies, quantisation trade-offs, GPU memory planning.

  • VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker); confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).

  • Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.

  • Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab’s platform team.

  • Peer-to-peer or distributed-systems background.

  • Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.

Important information for candidates

Recruitment scams have become increasingly common. To protect yourself, please keep the following in mind when applying for roles:

  • Apply only through our official channels. We do not use third-party platforms or agencies for recruitment unless clearly stated. All open roles are listed on our official careers page: https://tether.recruitee.com/

  • Verify the recruiter’s identity. All our recruiters have verified LinkedIn profiles. If you’re unsure, you can confirm their identity by checking their profile or contacting us through our website.

  • Be cautious of unusual communication methods. We do not conduct interviews over WhatsApp, Telegram, or SMS. All communication is done through official company emails and platforms.

  • Double-check email addresses. All communication from us will come from emails ending in @ tether.to or @ tether.io

  • We will never request payment or financial details. If someone asks for personal financial information or payment at any point during the hiring process, it is a scam. Please report it immediately.

When in doubt, feel free to reach out through our official website.

Read the full description
Finance REMOTE - Financial controller

Oversees financial operations, accounting systems, and reporting as the primary financial officer for the organization.

Lead Remote Posted about 24 hours ago Himalayas
What this role involves
About Dalstrong“There are NO Limits” is our motto.
Read the full description
Engineer Staff Software Engineer, AI-Native Systems

Designs and builds backend infrastructure and systems for an AI-native virtual care platform as a technical lead.

Lead Remote Posted 1 day ago Jobicy AI
What this role involves
Staff Software Engineer, AI-Native Systems (Tech Lead) Location: Remote (US) We are Virtual care only works if the infrastructure behind it does. Wheel builds that infrastructure — the systems that...
Read the full description
Sales OpenBase LLC: Head of Growth & Operations

Owns go-to-market strategy, customer acquisition, developer activation, and early-stage business execution to grow an AI gateway platform to 10,000 users.

Lead Remote Posted 1 day ago We Work Remotely — Programming
What this role involves

Headquarters: US
URL: https://openbase.ai/

Head of Growth & Operations — Openbase.ai

Full-time | Remote | Reports directly to the Founder

About Openbase

Openbase.ai is an AI gateway that helps developers and businesses access multiple AI models through one API. We simplify model access, routing, and usage management so teams can focus on building AI products.

Our MVP is ready. We are hiring a hands-on leader to own our go-to-market strategy, day-to-day business execution, and journey from launch to our first 10,000 users.

The opportunity

You will work directly with the founder to turn Openbase into a growing business with repeatable customer acquisition, strong retention, and disciplined operations.

You will own the commercial execution: finding our first customers, testing acquisition channels, improving onboarding, building partnerships, and organizing the people and systems needed to grow.

Our ambition is 10,000 users. Your responsibility is to build a credible, measurable path toward that goal, with clear milestones for activated developers, retained customers, paying accounts, and profitable usage.

What you will own

1. Go-to-market and customer acquisition

  • Identify our strongest initial customer segments, including AI startups, independent developers, agencies, and software teams.
  • Develop our positioning and explain why developers should choose Openbase.
  • Launch and improve acquisition across X, SEO, visibility in AI search engines, developer communities, partnerships, and targeted outbound.
  • Test paid campaigns with clear budgets, conversion tracking, and criteria for scaling or stopping.
  • Build a growth roadmap with weekly experiments and monthly acquisition targets.

2. Developer activation and retention

  • Own the journey from website visit to signup, first successful API request, first payment, and ongoing usage.
  • Work with engineering to remove onboarding friction and improve documentation, quickstarts, and integration examples.
  • Speak directly with developers to understand why they try Openbase, keep using it, or leave.
  • Establish onboarding communications, customer support processes, and reactivation campaigns.
  • Turn successful customer implementations into case studies and referrals.

3. Sales and partnerships

  • Personally source and close early customers.
  • Build a qualified pipeline of startups and businesses with recurring AI usage.
  • Develop distribution partnerships with developer tools, AI communities, accelerators, and complementary platforms.
  • Coordinate technical evaluations and customer requirements with engineering.
  • Build a repeatable sales process as demand grows.

4. Business execution and team management

  • Translate company goals into weekly priorities, accountable owners, and delivery dates.
  • Recruit and manage freelancers, agencies, and future hires across growth, content, developer relations, and customer success.
  • Own marketing budget allocation and evaluate spending against business results.
  • Coordinate customer needs with product and engineering priorities.
  • Implement practical workflows and automation that reduce manual work.

5. Metrics and commercial performance

  • Maintain a dashboard covering acquisition, activation, retention, paying customers, and API usage.
  • Track customer acquisition cost, conversion rates, and retention by channel.
  • Separate customer spending from provider costs and Openbase’s retained revenue and gross profit.
  • Monitor promotional credits and infrastructure costs so growth remains economically sound.
  • Give the founder a clear weekly view of results, blockers, and decisions needed.

What we are looking for

  • Experience leading growth and commercial execution at an early-stage SaaS, API, developer-tools, or AI infrastructure company.
  • Evidence of acquiring users from an early starting point, with specific numbers and your personal contribution.
  • Understanding of developer audiences and how technical products earn adoption.
  • Ability to understand an API integration, follow documentation, and discuss onboarding issues with engineers.
  • Hands-on ability to run experiments, analyze funnels, conduct customer calls, and manage delivery.
  • Experience managing budgets and external specialists.
  • Strong written English, sound judgment, and comfort working directly with a founder.

Experience with usage-based pricing, developer relations, technical SEO, and AI model platforms is particularly relevant.

Why join

  • Direct ownership of Openbase’s launch and growth.
  • Close collaboration with the founder and engineering team.
  • The opportunity to build the commercial function from its earliest stage.
  • Scope to take on broader company leadership as the business grows.

How to apply

Send your CV or LinkedIn profile, together with brief answers to:

  1. Which product have you helped grow? Include starting users, ending users, timeframe, budget, and your exact role.
  2. Which acquisition channels produced retained or paying customers?
  3. How would you approach Openbase’s first 1,000 activated developers?
  4. What would you personally execute, and where would you use specialists?
  5. What are your location, availability, and compensation expectations?

Include an example of a growth experiment or operating system you built. Share only information you are authorized to disclose.

To apply: https://weworkremotely.com/remote-jobs/openbase-llc-head-of-growth-operations

Read the full description