Infrastructure Engineer, Platform and Nodes
EngineeringNew York CityFull time
Run the boxes, the databases and the chain nodes that a trading app cannot be slow or wrong without.
About the role
We run our own infrastructure on dedicated hardware across several regions, because latency to venues and control over data are product features. That includes self-hosted chain nodes (a Base full node on Reth, a Hyperliquid non-validator node next to the validators, a TON liteserver with a full indexing stack), Postgres with point-in-time recovery, Redis, a leased background-loop layer and blue-green deploys with no user-visible downtime.
This role owns all of it. You will be the person who knows why the node fell off tip, what the vacuum is doing, and which change goes out first.
What you will do
- Keep every node at tip and every service ready: readiness that means something, alerting that fires before users notice.
- Own Postgres: schema migrations as expand and contract, index design for real query shapes, vacuum and lock behaviour under load, backups you have actually restored.
- Run deploys and rollbacks: blue-green cutovers, leases so two instances never fire the same side effect, config through a secrets manager.
- Plan capacity: disk for chain data that grows every day, memory for indexers, network for market-data fan-out.
- Harden the edge: firewalling, private networking, TLS, edge caching for public reads, rate limits.
- Write runbooks and automate the boring parts so incidents get shorter every quarter.
What we are looking for
- Linux at depth: processes, memory, disks, networking, and the tools to see what a box is doing right now.
- Postgres at depth, including behaviour under high write volume and migrations with minimal downtime.
- Docker and Compose in production, with a clear view of when a scheduler would and would not help.
- Networking across the stack: HTTP, WebSockets, TLS, DNS, and what goes wrong at each layer.
- Scripting that other people can read (Bash, Python), and the judgment to know when a script should become a service.
- Incident experience: you have been paged, found the cause, and fixed the class, not just the instance.
Nice to have
- Running blockchain nodes (Reth, Geth, Solana validators or RPCs, Hyperliquid, TON) and keeping them synced.
- TimescaleDB or ClickHouse at terabyte scale.
- Rust, enough to read the services you operate and fix a bug in one.
- Kafka or Redpanda, Cloudflare, Tailscale, Ansible or Terraform.
We encourage you to apply even if you do not meet every line above. Strong people rarely tick every box.
How to apply
Email us at hi@tryalpha.xyz with “Infrastructure Engineer, Platform and Nodes” in the subject line, a link to work you are proud of, and a few lines on the hardest system you have shipped and what broke. No cover letter, no form. A real person reads every application.
