Skip to the content

Projects01 / 07 · proptech · USA · 2023–2026

myHomeIQ · Intelligence 01

The market-intelligence service behind a US proptech platform — it knows which agent sold what, where, with whom, and which deals got away.

Senior Backend Engineer · #2 of 14 contributors · author of the repository’s third pull request
Skip to the engineering
In plain words

A realtor’s whole business is people who forgot them

A US realtor lives on repeat business: the person you sold a house to will sell it again in five to seven years, and refinance twice in between. By then the client has forgotten you exist.

This service is the platform’s intelligence layer. It ingests every recorded sale and mortgage in the country over eight years of history and turns that into something an agent can act on: who is selling in their area, who they compete with, who they should partner with — and which of their own past clients just did a deal through somebody else.

Two facts make it hard. The data arrives dirty, from dozens of government registries, MLS databases and paid providers, in different shapes and with the same person spelled five ways. And every external lookup costs real money, so the system has to remember what it already knows.

For engineers

Entity resolution, ETL, and PostgreSQL that had to stop growing sideways

32 tables plus 80 partitions — 60 of them by year — 617 indexes, 62 documented API operations.

Entity resolution · the long war (35+ PRs)

Three matching strategies in priority order — licence number → name + phone → name + company — with BFS expansion of duplicate groups, batched upsert_all, a 128 MB cache, and resumable progress so a multi-hour run restarts where it stopped. Fuzzy name matching is raw SQL on pg_trgm: similarity ≥ 0.6 intersected with JSONB phone arrays via ?|.

Then the same problem, vectorised

My largest pull request here — 3 645 hand-written lines — replaced an LLM-per-request approach with 768-dimension embeddings on pgvector + HNSW, building a “profile sentence” per realtor out of names, companies, cities and ZIPs. Stated honestly: merged behind a feature flag, then reworked out and dropped from the schema in September 2025 — it never carried production load.

PostgreSQL under load

RANGE partitioning of the month-stats tables, plus the 2025 and 2026 year partitions on sales and loans — the original sales/loans conversion was a colleague’s. Read/write splitting in production traffic via connects_to: roles wired up, the database selector placed correctly in the middleware stack, and the writing role forced at the exact points where read-after-write broke. 3 619 lines across four merged commits.

A new data supplier, end to end

A pipeline built from nothing for a new market-data feed: S3 download → parse → validate → batched upsert → merge into existing sales, guarding against overwriting co-listing data. Three large PRs of 2.4–3.1k lines, plus a three-level fallback matcher for agent contacts.

Search that got smaller

OpenSearch filters, sorting and relevance for agent search. Reindexing moved off model callbacks onto a queue with a periodic job, and an OOM during bulk reindex fixed by changing mode and throttling. One search PR deleted 7 919 lines, of which 7 131 were recorded HTTP cassettes: simplification, not accumulation.

Scoring, and knowing when not to pay for it

External move-probability scoring behind my own Pipelines / Providers / Sources abstraction, backfilled in the background, then exposed as dashboard filters and sorting. Outside production the provider is swapped for a random one, so the paid quota is never burned by a test run.
272
pull requests merged
184
tickets closed end to end
91.9%
of my PRs touch specs
62
documented API operations
Scope, honestly. Helm charts, Dockerfiles and CI pipelines were owned by dedicated people on this platform — I worked inside a Kubernetes/CircleCI environment rather than building it. My contribution is applied backend and data work. Every number here is counted from git history, at my last commit in February 2026.

Who worked on it

Solo — the whole stack

Danyil Shkoropad · Senior Backend Engineer · 2023–2026
  • The service’s public API, Lost Deals, and three years of entity resolution.
  • Partitioning and read/write splitting as history reached eight years.
The other 13 contributors were the client’s own team — not people from this studio.
Open CV

Read to the end and want to talk it over?

rubyco.in