ByteCustodian

The whole life of your data, in one place.

What is ByteCustodian?

Functionality of four products packaged in a single product.

Lakehouse

Open Parquet files you own, on commodity storage, tracked by a catalog in your own database. Nothing locked in a proprietary store.

Analysis engine

One machine, right-sized to the data you actually have. No cluster to size, no shuffle between machines.

Reporting tool

The analysis renders its own dashboard, where the data already sits. Nothing to buy, wire up or charge per person.

Governance platform

Masking and row rules enforced as the query runs, and every result traceable to its versioned inputs.

One install. One set of rules. One bill.

ByteCustodian in one minute

Request a demo

A demo shows you the product working. A proof of concept runs it on your own data, free: your real data, an anonymized copy of it, or synthetic data we generate to match your volumes. You see real numbers at your scale before you spend anything.

To ask for either, tell us whether you want a demo or a proof of concept, roughly how much data you have, and whether it should run in your AWS account or on your own hardware. A person replies. Write to us: info@subtractsoftware.com

Engineered from first principles.

Our product

Our principles

How do you control access to sensitive data?

One column, viewed by three different personas.

One customer record, as three people see it
Column An analyst sees Support sees Compliance sees
SSN ●●●-●●-3421 123-45-6789
Email a•••@example.com alice@example.com
Card number ●●●● ●●●● ●●●● 4242

Rows are limited the same way. A regional analyst sees only the rows for their own region.

The rule is applied when the query runs, so there is no masked copy to build and keep in sync. How it is enforced, signed in to and tested, for your IT and risk team.

Why keep everything in one place?

No extracts. No second copies.

Most data platforms make you move your data to get value from it — and you lose control of it in the process. ByteCustodian keeps the entire lifecycle in one governed place. Data is ingested, stored, queried, analyzed, and turned into a dashboard without ever leaving the lakehouse. No downstream tool has to be told the rules again.

Your data today

  • Warehouse peak-priced
  • Data lake stale, ungoverned
  • Exports files sent by email
  • Spreadsheets on someone's laptop
  • BI tool billed per seat
  • Compliance bolted on after

In ByteCustodian

The lakehouse
  1. Ingested
  2. Stored
  3. Queried
  4. Analyzed
  5. Dashboard

One catalog. One access model. One bill.

Where does it run?

Two configurations, both inside your boundary.

  • In your own AWS account.
  • On your own hardware, with S3-compatible storage on your own disks.

Both are single-tenant. Either way, you hold the data and the format, and it never leaves your AWS account or your servers.

Which personas does ByteCustodian serve?

If your business runs on data, ByteCustodian is for you.

You get a right-sized strategy for that data and the platform to carry it out, as one whole — storage, analytics, reporting and governance together — without a bill that spikes or a cluster you don't need.

  • The analyst

    “Can I answer a business question today?”

    Browse the datasets, query them, run the heavy analysis, and publish the dashboard. No waiting on another team. No manual.

  • The compliance officer

    “Where is our sensitive data, and who can see it?”

    Scans find it. Masks and row rules protect it. Every decision and every query is on the record.

  • The administrator

    “Who has access, what is running, and what will it cost?”

    One access model. A live view of every query. Upkeep that runs itself. A spending ceiling you set.

Ask for a demo and we will show you the product through each of these three pairs of eyes: info@subtractsoftware.com

Do you have any measured benchmarks?

Nothing was warmed, cached, or repeated until it looked good.

Results we measured on a billion-row banking book
The work, on a billion-row banking book Measured
A billion governed rows, aggregated three ways38 seconds
A billion rows joined to a billion rows82 seconds
One analyst's query, limited to the rows she is allowed to seeabout 1 second
A full day of random work on a desktop built in 20121,119 analyses, 0 restarts
What we charge per query, per user, or per byte$0
How these were measured

Every run went through the governed engine, with row filters, column masks and access checks in force. The first three ran cold on one 16-core AWS machine, started for the run and shut down after it. The fourth ran on an office desktop from 2012, on purpose: eight cores, 16 GB of memory and one ordinary disk, carrying the platform, the storage and the database on its own. Any server you would install this on today is several times more powerful. We wanted to see the platform on hardware nobody would choose.

  • The benchmark paper

    Two pages, and one page of evidence. The big work timed on a billion rows, and the monthly bill worked out at the cloud provider's own list prices.

    Ask for the benchmark paper: info@subtractsoftware.com

  • The endurance paper

    Two pages. Twenty-four hours of analyses, compliance scans and maintenance, fed at random to a desktop from 2012, on purpose — and what held.

    Ask for the endurance paper: info@subtractsoftware.com

What does it cost?

Our price never follows how much you use it.

Nothing per query, per user, per byte or per dashboard. A metered platform charges you more for working harder, so people start rationing — they ask the smaller question and skip the check that would have been interesting. We think that is backwards.

Your whole cost is two lines. Our licence, paid to us: one flat number a year, which never moves with use. The machines, paid straight to AWS at AWS’s own prices: a base that runs all month, what you store, and any large machine a heavy job rents. That second line is the only one that moves. It is AWS’s and not ours, and you set its ceiling. Here is the base in full.

The base cluster's monthly cost in your own AWS account, at AWS list prices
The base that runs all month Per month
Small — everyday work for a small team2 × t3.large · RDS db.t4g.micro · EKS control plane · NAT gateway · load balancer about $280
Medium — twice the standing capacity2 × t3.xlarge · RDS db.t4g.small · EKS control plane · NAT gateway · load balancer about $410
Large — four times the standing capacity2 × t3.2xlarge · RDS db.t4g.medium · EKS control plane · NAT gateway · load balancer about $680

AWS on-demand list prices, us-east-1, as of September 2026. AWS sets these prices and can change them.

Nothing runs while nobody is working. Most data platforms keep large machines running around the clock, and bill you for every idle hour. We think that is waste — of your money, and of the energy those machines burn doing nothing.

ByteCustodian keeps that small base running, and that is all. Heavy work starts a large machine for that one job. The machine shuts down within minutes of finishing, and AWS stops billing for it. You set the ceiling, and nothing can spend past it. On your own hardware there is nothing to rent: heavy work waits its turn, and the machine you already own does it. A smaller bill, and a smaller footprint.

Our licence is the larger of the two lines. We say that plainly so the figures above are not mistaken for the price. It is one flat number a year, set by the size of your institution and never by how much you use the platform.

What the base covers, and what it does not

None of these is a capacity limit. The base is sized to be cheap, not fast. They are what stays up all month: the machines that run everyday analysis, the small database that holds the catalog, and the networking, which includes the load balancer in front of the platform and its public addresses. Heavy work does not run here — it rents its own machine, which is priced under “What heavy work costs”, below. Your data is a separate line, and it has no ceiling: AWS charges about $0.023 a gigabyte a month for S3, so a terabyte is about $23 a month and ten terabytes about $230, whichever base you run.

Hundreds of dollars, not thousands. Our analysis engine runs on one machine sized to the data you really have, so there is no cluster idling between questions. The small row is the base our AWS benchmark ran on; the billion-row work above rented its large machine from it. The other two are worked out from the same price list. On your own hardware there is no AWS line at all.

How the licence is set

Add analysts, add dashboards, run it overnight: the number does not move. What it should be for you depends on your size, your data and what it is replacing, and that is a short conversation rather than a guess. Tell us roughly what you are working with and we will put a real number in front of you: info@subtractsoftware.com

What heavy work costs

We do not assume your work is small. A job declares what it needs, and the platform buys the cheapest machine that fits it, runs that job on it alone, and shuts the machine down within minutes of finishing. AWS bills that machine by the second, with a one-minute minimum, so you pay for the minutes you actually use. Our benchmark machine — an m5.4xlarge, 16 cores, 64 GB, a 120 GB gp3 disk — is about $0.92 an hour, and most heavy jobs are minutes, not hours: the billion-to-billion join above ran in 82 seconds. The machine it rented existed for about three and a half minutes in all: starting up, running the join, and shutting down. At $0.92 an hour, billed by the second, that came to about five cents. Keeping that machine busy a full eight hours every working day would be an extreme, and even then it is about $160 a month. It is an example, not a ceiling. The machines on the menu are yours to set, and a job that genuinely needs something far larger gets something far larger. What you cap is how much may run at once.

Dashboards are included, with no charge per person. They run where the data already sits: one copy, one set of rules, nothing to export and nothing to reconcile.

Work it as hard as you like. That was always the idea.

Every figure on this page is representative, not a quote. They are real numbers from real runs and AWS’s own price list, as of September 2026, and they are the honest shape of what this costs. Your AWS line will differ with three things: how much data you store, which base you run, and how much heavy work you rent machines for. Our licence does not move with any of them.

We would rather earn your trust than win a line item — customer love is the only currency we keep score in.

  • What it would cost you

    Tell us roughly how much data you have, how many people will use it, and whether it should run in your AWS account or on your own hardware. We come back with a short conversation followed by real numbers for your situation.

    Start there: info@subtractsoftware.com

How long does it take to go live?

If software needs an army to go live, it was not designed well enough.

That is the standard we hold ourselves to, and it is the reason this company exists. We are Subtract Software. Taking things out is the work, and a system with less in it has less to configure and less to go wrong.

So there is no implementation to buy. No partner to engage, no statement of work, and nobody billing you by the hour while you wait to see your own data.

See it running today. One command on one Linux machine installs everything — the platform, the storage, the catalog, sign-in, certificates, sample people and sample data. It prints an address when it finishes.

Free onboarding hours. Installing into your own AWS account is real infrastructure work, and we will not pretend it is a single click. So we do it with you. You get 40 free hours of our time to install it, connect your first sources, set up your groups and masking rules, and build your first reports. Free. Not billed, and not added to the invoice later. We expect that to be enough for most institutions to be live in days, not quarters.

If your data is hard to extract

A dozen source systems, or decades of history, is a real conversation, and we will be straight with you about what it takes. We expect most institutions never face this. We would rather spend those hours learning your business than billing you for ours.

Your administrators will thank you. The lake stays fast and lean on its own — compaction, snapshot expiry, and catalog housekeeping run on a schedule.

Simple to install. Simple to operate. No babysitting.

See it on your own data.

We will run a free proof of concept with you. Use your own data, an anonymized copy of it, or synthetic data we generate to match your volumes. You see real numbers at your scale before you spend anything.

Request a demo, pricing, or a proof of concept: info@subtractsoftware.com