!

When none of the 94 pre-built datasets is the one you need

Photo of Alex Wilkinson
Alex Wilkinson
CEO of Houski
2026-07-13

The datasets page has around 94 pre-built cuts across 15 categories. Flood risk properties. Empty nesters and retirees. Roofs needing replacement. Punjabi and South Asian communities. Heat pump conversion targets. Vacant land. Cul-de-sac homes.

They cover a lot. They will not cover you, because your business is not the average of 94 businesses.

So the sixteenth category is the one nobody markets: we build the dataset you actually need.

What a custom dataset is

Our data team assembles a dataset to your specification - combining data points across the schema, applying whatever filters your business runs on, and developing analytics specific to what you are deciding. It gets delivered in your format, comma-separated values or JSON, on your schedule. One-time, monthly, or whatever cadence your work runs at.

That is the whole offer. It is deliberately unglamorous.

When to ask for one

Your filter does not fit in a query string. Some criteria are a rule, not a filter. "Houses over 1,600 square feet, built before 1995, in areas where the median after-tax income of residents clears 70,000, but excluding anything within two blocks of the addresses we already serve, weighted by how many of our existing customers live on the street." That is a program. You can write it against the API, or we can hand you the answer.

Your boundary is not a city. Cities are a convenience, not a business. Your service area is a drive-time polygon, a franchise territory, a set of postal walks, a sales rep's patch. Hand us the boundary, get back the properties inside it.

You want derived fields we do not publish. Roof age in years rather than install year. A composite fit score built from your own weights. A rank within neighbourhood rather than a raw value. Distance to your nearest depot. These are trivial for us and a project for you.

You need it suppressed against your own list. The most valuable filter is often the one that removes people. Existing customers, addresses you mailed last quarter, leads that already said no. Send us the exclusion list, get back only the addresses worth spending on.

You need it on a schedule that matches your operations. A monthly mail drop wants a monthly file. A quarterly planning cycle wants a quarterly file. Not a one-time export you re-buy when someone notices it is old.

You want a shape a query does not produce. Rolled up to neighbourhood with the counts and medians already computed. Pivoted for a dashboard. Split by territory into one file per rep.

Why not just query the API

Sometimes you should. If you have engineers, an API key, and a clear filter, the properties endpoint will do this and you will not need us.

Custom datasets are for the case where writing that is not a good use of your team, or where you do not have that team at all. Plenty of our customers are marketers, underwriters, planners, and operators rather than developers. The pitch is not that the API is hard. It is that you should not have to care.

The other case is scale. Pulling a few thousand properties is an afternoon. Pulling several million with a dozen joins and a scoring pass is a pipeline, and pipelines want owning. If the alternative to a custom dataset is a project plan, ask for the dataset.

The freshness part, which is the part that matters

A custom dataset is not a one-time export with your name on it.

It gets rebuilt on your schedule, against data that is updated daily. New construction shows up. Sold and expired listings drop out. Assessments reroll. Estimates improve. Boundaries move and the joins get rerun to match.

That is the difference between a custom dataset and a list you buy. A list is correct on the day it is cut and rots silently from then on, and nothing about the file will ever tell you. A rebuilt dataset is correct every time you open it.

If you have ever bought a marketing list and watched the response rate sag over a year, that is what you were watching.

What to send us

The request goes better when you bring:

  • What you are deciding. Not the fields you think you want, the decision you are making. We know the schema. You know the business. The fields are our job.
  • Your geography. City list, boundary file, postal walks, coordinates. Whatever you have.
  • Your exclusions. Existing customers, prior campaigns, anything already touched.
  • Your format and cadence. Comma-separated values or JSON, and how often.
  • Your volume. Ten thousand rows and ten million rows are different conversations.

You do not need a specification. A paragraph describing the job usually gets further than a field list, because half the time the field list is wrong in ways we can see and you cannot.

Start with the pre-built ones anyway

Even if you know you need something custom, browse the datasets page first. Not because one will fit, but because each one shows its filters and its example use case, and it is the fastest way to see what is possible. Most custom requests we get are a pre-built dataset with two changes. Finding the one that is close makes the conversation short.

Getting started

  1. Browse the datasets and find the closest thing to what you want
  2. Request a custom dataset and describe the decision, not the schema
  3. We scope it - fields, filters, geography, format, cadence
  4. You get a sample before you commit to the full build
  5. It rebuilds on your schedule, against data that moved since last time

Ninety-four datasets is a lot of guesses about what people need. If none of them guessed you, that is not a gap in the catalogue, it is the reason the data team exists. Ask for the one you need.