Comparable-sales valuation
Serenada West Sec 2 · Georgetown, Texas 78628 · Williamson County · MLS
The property
Land is priced at this parcel’s own Williamson CAD rate — $192,386 across 44,431 square feet, or $4.33 per square foot — not a county average and not a third-party record. The county roll reconciles for every year on file: land plus improvement equals total.
Limits
This is a comparable-sales opinion of value, not a certified appraisal, and it is not a mortgage or lending document. It is produced from the ABOR feed, Williamson CAD records and ALEX's own adjustment engine, run automatically with no comparable selected by hand. Six comparables support this figure; how firmly they agree is measured, not assumed, in the sections below.
Competition
Four properties. Reference only — they are not used
in the valuation and carry no weight in it. Every figure in this report comes from
closed sales — properties where money actually changed hands. The listings below
have not sold. An asking price is a seller’s opinion of value, not a measurement of it, and no
asking price enters this calculation at any point.
They are shown because they are useful context for a decision, not for a valuation: a buyer
standing in this house can see these the same afternoon, and they are the alternatives
this property competes against. All appear on the
map below as blue diamonds on dashed lines.
| Listing | Distance | Asking | Living area | $/sf | Built | Days on market |
|---|---|---|---|---|---|---|
| 4301 Verde Visthis property | — | not listed | 2,052 sf | 1982 | — | |
| 300 W Sequoia Spurnorth-north-east | 0.29 mi | $2,700 | 2,496 sf | $1.08 | 1979 | 0 |
| 4319 Miramar Drwest | 0.29 mi | $540,000 | 2,043 sf | $264.32 | 1993 | 41 |
| 205 E Esparada Dreast | 0.80 mi | $675,000 | 2,500 sf | $270.00 | 1974 | 83 |
| 1011 Serenada Dreast-north-east | 1.84 mi | $424,000 | 1,846 sf | $229.69 | 1983 | 69 |
300 W Sequoia Spur
$2,700 · 2,496 sf · built 1979 · 0.29 mi NNE · 0 days on market
It offers 444 more square feet — built 3 years earlier. On the market 0 days without selling.
4319 Miramar Dr
$540,000 · 2,043 sf · built 1993 · 0.29 mi W · 41 days on market
It offers 9 fewer square feet — built 11 years later. On the market 41 days without selling.
205 E Esparada Dr
$675,000 · 2,500 sf · built 1974 · 0.80 mi E · 83 days on market
It offers 448 more square feet — built 8 years earlier. On the market 83 days without selling.
1011 Serenada Dr
$424,000 · 1,846 sf · built 1983 · 1.84 mi ENE · 69 days on market
It offers 206 fewer square feet — built 1 year later. On the market 69 days without selling.
To be explicit, because it matters: nothing on this page was adjusted, weighted or influenced by any of these listings. They appear in this section and on the map, and nowhere else in the calculation. If all were withdrawn tomorrow the valuation would be unchanged to the dollar. What they tell you is what this property is being shopped against — which is a question about strategy, not about value.
What these numbers are
$502,747 is what we believe the seller keeps at closing — not what the contract would read. It is the value after the seller has paid the buyer’s closing costs and made the repairs the buyer asked for. Every comparable sale below is stated the same way, so like is compared with like.
Both deductions, from this report’s own evidence
| 303 Sequoia Spur — a closing-cost concession | |
| Closed price | $550,000 |
| Buyer closing costs paid by the seller | −$300 |
| Repairs at the buyer’s request | none |
| Net sale price used here | $549,700 |
| 4300 Madrid Dr — a closing-cost concession | |
| Closed price | $442,500 |
| Buyer closing costs paid by the seller | −$23,000 |
| Repairs at the buyer’s request | none |
| Net sale price used here | $419,500 |
609 Esparada Dr and 506 La Paloma Dr had neither deduction, so for each the net and closed price are the same. Two houses can close at the same contract price and not be the same sale: if one seller paid $12,000 toward the buyer’s costs and the other paid nothing, the first achieved $12,000 less. Working in net sale price removes that difference before anything is compared, which is why these comparables are directly comparable to one another and to this property.
Where everything is
Everything is within about 2 miles, and every marker is clickable — a red circle jumps to that sale in the table below, a blue diamond to the for-sale property in its own section. Solid red lines join the six comparables to the subject; dashed blue lines join the four properties currently for sale, which are shown for reference and are not used in the valuation.
Comparable sales — used in the valuation. 4300 Madrid Dr 0.12 mi north, 506 La Paloma Dr 0.21 mi south-south-east, 303 Sequoia Spur 0.22 mi north-north-east, 609 Esparada Dr 0.24 mi west-south-west, 126 Esparada Dr 0.64 mi east, and 806 Serenada Dr 1.64 mi east-north-east. All six sold, all six within about 2 miles. They are joined to the subject by solid red lines, and each one carries the adjustments set out later in this report.
Properties for sale — reference only, not used in the valuation. 300 W Sequoia Spur 0.29 mi north-north-east, 4319 Miramar Dr 0.29 mi west, 205 E Esparada Dr 0.80 mi east, and 1011 Serenada Dr 1.84 mi east-north-east. None has sold, so none is evidence of value; they are joined to the subject by dashed blue lines to mark that difference. All are shown because a buyer viewing this property will also be viewing them: see Other Properties for Sale.
Not Market Averages
Nothing here is a market average. Every figure is computed from this property’s own six comparable sales — how far the method missed when each was hidden and re-predicted, how much the six disagree with one another, how far they had to be moved to match this house, and how old they are. Open either panel below: the first states each measure in plain terms, the second gives the formal treatment.
1. Leave-one-out back-test (mean absolute percentage error). For i = 1..n the full pipeline — comparable adjustment across all 4 factors, then adjustment-weighted reconciliation — is refit on D \ {(xᵢ, yᵢ)} and evaluated at xᵢ. The reported statistic is n⁻¹ Σ |ŷ₋ᵢ − yᵢ| / yᵢ. LOOCV is approximately unbiased for expected out-of-sample risk but has high variance: the n training sets share n−2 observations, so fold estimates are strongly positively correlated and the apparent precision is optimistic (Hastie, Tibshirani & Friedman, ESL 2nd ed., §7.10). At n = 6 this is the binding caveat and the reason the figure is presented as descriptive of this comparable set rather than as a population claim.
2. Adjusted spread — a dispersion measure, not an error measure. (max vᵢ − min vᵢ) / V̂ over the adjusted comparable values. It quantifies agreement among estimates, which is not the same quantity as distance from the truth and can move in the opposite direction. Comparables adjusted by one rate table share method and data; their errors are positively correlated, so a tight spread partly reflects common specification rather than convergence on the right answer. Formally, for estimates with pairwise error correlation ρ, the variance of their mean is (σ²/n)[1 + (n−1)ρ] — under strong positive ρ the effective sample size is far below n, and observed spread understates true uncertainty. Low spread with a shared bias is the classic route to a confident wrong answer, which is why this page reports spread and held-out error separately and never substitutes one for the other.
3. Mean absolute adjustment — a proxy for accumulated model error. n⁻¹ Σ gᵢ/Pᵢ, gross adjustment over net sale price. Every adjustment is itself an estimate carrying its own error, so total adjustment magnitude bounds how much of a final figure is model rather than observed transaction. Under the approximation that adjustment errors are independent with variance proportional to magnitude, the error contributed grows with √Σg² — which is the premise the reconciliation weighting rests on, and the reason gross rather than net is the right summary. It is an ordering statistic, not a calibrated error: it ranks comparables by how much work was done to them without asserting how wrong that work was.
4. Mean comparable age — exposure to temporal drift. Mean days between each comparable’s close date and the valuation date. Time adjustment removes the expected market movement, so residual risk is the variance of the index estimate over the carry interval, not the movement itself. Under the Case–Shiller error decomposition that variance grows roughly linearly in the interval, Var ≈ σ²ᵘΔt + 2σ²ᵤ, so age is a proxy for how much index uncertainty each comparable imports. Here the mean is 120 days; the relevant exposure is not the elapsed time but the month-to-month volatility of the index across it, which the market-timing section quantifies directly.
What the four are jointly evidence of. Taken together they describe the provenance of the estimate, not its accuracy: how much of the final figure is observed transaction and how much is model (mean absolute adjustment), whether the independent indications agree (adjusted spread), how much temporal extrapolation was required (comparable age), and how the method performed when it could not see the answer (the back-test). Only the last is an error measure. The first three can all look excellent on a valuation that is wrong.
Why they cannot simply be combined. A reader’s instinct is to average them into a single quality number. They are not independent, and the dependence runs the wrong way. Comparables selected for similarity have small adjustments by construction — the cascade chose them for it — so a low mean adjustment partly reflects the selection rule rather than the market. Tight agreement among comparables adjusted by one rate table reflects the common rate table. Formally, for estimates with pairwise error correlation ρ, the variance of their mean is (σ²/n)[1 + (n−1)ρ]: at ρ near 1 the effective sample size approaches 1 regardless of n, and observed spread understates true uncertainty by an unknown factor. Nothing here estimates ρ — six observations cannot.
The failure mode this section exists to expose. Comparables that are genuinely alike, adjusted by a shared and slightly wrong rate table, will produce a tight spread, small adjustments, and a low back-test error — while every one of them is biased in the same direction. Correlated error is invisible to all four measures simultaneously.
What none of these four is. None is a confidence interval, and none should be read as one. Coverage statements live in Cross Validation, where they are derived from the held-out error distribution rather than from agreement among comparables — agreement among correlated estimates being precisely the thing that can be high while the answer is wrong.
The land
The comparables sit on 0.51 to 1.47 acres; this property sits on 1.02. Any difference is priced at a rate taken from the target’s own parcel — what Williamson County says its land is worth, divided by how much land the county says it has. Nothing here is estimated or taken from a listing portal.
| Parcel | CAD acres | Lot size | CAD land value | County $/sf |
|---|---|---|---|---|
| 4301 Verde Visthe subject, R045830 | 1.020 | 44,431 sf | $192,386 | $4.33 |
| 609 Esparada Dr | 0.510 | 22,216 sf | $120,938 | $5.44 |
| 303 Sequoia Spur | 1.010 | 43,996 sf | $191,196 | $4.35 |
| 4300 Madrid Dr | 1.020 | 44,431 sf | $192,386 | $4.33 |
| 126 Esparada Dr | 0.840 | 36,590 sf | $169,852 | $4.64 |
| 806 Serenada Dr | 0.550 | 23,958 sf | $127,517 | $5.32 |
| 506 La Paloma Dr | 1.470 | 64,033 sf | $239,657 | $3.74 |
The county’s acreage and the MLS lot size agree exactly here — 1.020 acres is 44,431.2 square feet, and the listing says 44,431.2 — so the two independent records of this parcel’s size do not conflict.
| Comparable | Lot size | Difference | Land adjustment | Share of net |
|---|---|---|---|---|
| 609 Esparada Dr | 22,216 sf | −22,216 sf | +$96,193 | 19.9% |
| 303 Sequoia Spur | 43,996 sf | −436 sf | +$1,886 | 0.3% |
| 4300 Madrid Dr | 44,431 sf | 0 sf | $0 | 0.0% |
| 126 Esparada Dr | 36,590 sf | −7,841 sf | +$33,950 | 6.1% |
| 806 Serenada Dr | 23,958 sf | −20,473 sf | +$88,648 | 18.2% |
| 506 La Paloma Dr | 64,033 sf | +19,602 sf | −$84,876 | 11.8% |
Market timing
A sale that closed months ago is a snapshot in time. To use it as evidence about today, it has to be moved forward by however much the market moved in between. That is what the market timing line on each comparable does.
| Comparable | Closed | Index then | Index now | Factor | Adjustment |
|---|---|---|---|---|---|
| 609 Esparada Dr | 2026-01 | 166.975 | 167.172 | 1.0012 | +$568 |
| 506 La Paloma Dr | 2026-04 | 171.921 | 167.172 | 0.9724 | −$19,889 |
| 4300 Madrid Dr | 2026-05 | 166.970 | 167.172 | 1.0012 | +$506 |
| 303 Sequoia Spur | 2026-06 | 171.992 | 167.172 | 0.9720 | −$15,406 |
| 806 Serenada Dr | 2026-06 | 171.992 | 167.172 | 0.9720 | −$13,674 |
| 126 Esparada Dr | 2026-08 | 165.128 | 167.172 | 1.0124 | +$6,868 |
Adjustment = net sale price × (factor − 1). All come from the Williamson County index, built from 19,824 repeat-sale pairs.
What the index numbers mean. They are not dollars and not percentages — they are a scale, and every scale needs a zero. Here the zero is January 2016, set to 100. That month is the base by construction: the regression fixes its value and measures every later month against it. So a reading of 167.172 for September 2026 means Williamson County house prices are 67.2% above where they stood in January 2016.
Specification. Bailey, Muth & Nourse (1963), A Regression Method for Real Estate Price Index Construction, JASA 58(304), 933–942. For a dwelling transacting at periods s and t with s < t, the model is ln(Pᵗ/Pₛ) = βᵗ − βₛ + εₛᵗ. Let X be the N × T design matrix with xᵢₛ = −1, xᵢᵗ = +1 and zero elsewhere. Then β̂ = (X′X)⁻¹X′y with yᵢ = ln(Pᵗ/Pₛ). The base period is normalised to β₀ = 0 and its column deleted; retaining it makes X′X singular, since the indicators sum to zero by construction. The index is Iₜ = 100·exp(β̂ₜ).
Identification. The estimator is consistent for the market factor under the assumption that E[ε | X] = 0 — that idiosyncratic price changes are mean-independent of when a property happened to transact. Dwelling fixed effects difference out exactly because the same unit appears with opposite signs on both sides, which is what removes omitted-variable bias from unobserved quality and is the method’s central advantage over a hedonic index.
Error structure. Case & Shiller (1987, New England Economic Review; 1989, AER 79(1), 125–137) decompose εₛᵗ = (Hᵗ − Hₛ) + (uᵗ − uₛ), where H is a Gaussian random walk in the property’s own value and u is i.i.d. transaction noise. The implied variance Var(ε) = σ²ᵘ(t−s) + 2σ²ᵤ grows linearly in the holding interval, motivating their three-stage weighted least squares: OLS, regress squared residuals on the interval, then GLS with weights 1/√(â + b̂(t−s)). This index does not apply that weighting. It instead excludes holds under 180 days and annualised changes beyond ±60% outright — a trimming rather than a down-weighting strategy, which is more conservative, discards information the WLS would retain, and is stated rather than left implicit.
Shrinkage. County-month estimates are shrunk toward the metropolitan estimate in log space, ln Iᶜₜ = w ln Iᶜₜᶜ + (1−w) ln Iᵐₜ with w = n/(n+k), k = 20. This is the James–Stein / empirical-Bayes posterior mean under a normal-normal hierarchy where k is the ratio of within-month sampling variance to between-county prior variance; here it is fixed rather than estimated. The effect is that thin months borrow strength in proportion to how little evidence they carry of their own — September 2026, with 38 pairs, sits at w = 0.655.
Known limitations. (i) Sample selection. The index is estimated only on dwellings that transacted at least twice, which is not a random sample of the stock; properties that turn over frequently differ systematically, and no Heckman-type correction is applied. (ii) Revision. Estimates for recent periods move as later transactions arrive, so the most recent months are the least stable — a property shared with all repeat-sales indices including S&P CoreLogic Case–Shiller. (iii) Renovation. The dwelling-constant assumption fails where a property was materially improved between sales; the ±60% annualised filter removes the extreme cases and cannot remove the moderate ones. (iv) Aggregation. A single county index assumes a common market factor across submarkets within it.
Empirical note on smoothing. A trailing three-month average of the log index was tested against the raw monthly series by ten-fold cross-validation on 20,631 Williamson County pairs (September 2026), scoring out-of-sample RMSE of predicted log price change. It was worse in 10 of 10 folds (mean difference +0.00171, 95% CI [+0.00139, +0.00202], t = +10.70); a five-month window was worse again (t = +22.77). Monotonicity in the smoothing window is the signature of signal removal rather than noise suppression, consistent with the shrinkage above already performing the variance reduction a second pass would duplicate.
Cross Validation
How this was tested, in one sentence: we covered up one of the six sales, worked out what this method would have said that house was worth using only the other five, then uncovered it and measured how close we got — and did that six times, once for each sale.
| Held out | Actual net | Predicted from the others | Error |
|---|---|---|---|
| 609 Esparada Dr | $482,500 | $412,695 | −14.47% |
| 303 Sequoia Spur | $549,700 | $525,514 | −4.40% |
| 4300 Madrid Dr | $419,500 | $545,005 | +29.92% |
| 126 Esparada Dr | $554,955 | $529,318 | −4.62% |
| 806 Serenada Dr | $487,900 | $243,946 | −50.00% |
| 506 La Paloma Dr | $720,000 | $833,043 | +15.70% |
Part one — what a valuation should be measured against
This section is reserved. The standard is being supplied separately and will be placed here in full. It is deliberately empty rather than filled with a plausible-looking threshold: a benchmark that this report happened to pass would be indistinguishable, to a reader, from one it had actually been measured against. Until the paper is here, no accuracy standard is being claimed or implied anywhere in this report.
What is known today, pending that standard. Three sets of numbers exist and none of them is a standard — they are measurements, and they are not measuring the same thing:
1. This property. Median error 15.08%, FSD 33.21%, RMSLE 0.3159, 68% interval ±21.39%, from six held-out comparables. Specific to this house, and thin: six observations describe this comparable set accurately and generalise weakly.
2. This model, across its blind holdout. MdAPE 7.77%, MAPE 11.22%, RMSLE 0.1605 over 19,486 sales the model never saw during training. This is the right basis for any general claim about the method, and it is deliberately not what the tiles above report — those answer “how accurate is this valuation”, which is a different question.
3. What the vendors publish for themselves. HouseCanary: national MdAPE 2.8% over 1,994,203 transactions internally, 2.9% on a blind third-party test. Zillow: median error 1.79% on-market and 7.20% off-market. Two cautions before either is treated as a bar to clear. Both are national aggregates dominated by dense, homogeneous tract housing where an automated model does best, and neither is conditioned on the kind of property this is. And the on-market figure is not independent of the asking price — both vendors consume list price, which is the whole point made in the Benchmarks section.
What a standard would need to specify to be usable here, and what the forthcoming paper will presumably settle: which statistic is the criterion (MdAPE, FSD, coverage, or a hit rate); whether it is a point threshold or a distribution; whether it applies per property or per portfolio; how it is conditioned on price band, property type and evidence depth; and whether an interval must be calibrated — that is, whether a stated 68% interval must actually contain the outcome 68% of the time, which is a materially stricter requirement than a median error threshold and the one most AVM disclosures avoid.
Where this report already exceeds common practice, whatever the standard turns out to be: the comparables are named, every adjustment is itemised with its reason, the estimator and its weights are stated, and the held-out test results are printed so every summary statistic can be recomputed by hand. A standard can be applied to a number; it can only be verified against a number whose derivation is visible.
Part two — how this valuation was measured
Why hold anything out at all. Scoring a method on the same sales it was given measures how well it can interpolate data it has already seen, which is not the quantity anyone cares about. The quantity of interest is expected loss on an observation drawn from the same population and not used in fitting: Err = E(X,Y)[L(Y, f̂(X))]. Resubstitution error is downward-biased for it by construction.
Design: leave-one-out cross-validation. For i = 1,…,n the entire pipeline — all 4 adjustment factors, the repeat-sales time carry, and the inverse-gross-adjustment reconciliation — is refit on D−i = D \ {(xᵢ,yᵢ)} and evaluated at xᵢ, giving CV = n⁻¹ Σᵢ L(yᵢ, f̂−i(xᵢ)). Critically, the held-out comparable is excluded from the reconciliation as well as from the adjustment grid; leaving it in the weighting while removing it from the grid would leak the answer through the weights.
Properties, and the one that governs here. LOOCV is approximately unbiased for Err — it trains on n−1 of n points, so the pessimistic bias of k-fold at small k is largely absent. It pays for this in variance: the n training sets are nearly identical, sharing n−2 observations, so the fold-level estimates are strongly positively correlated and Var(CV) does not contract at the 1/n rate an independent average would. At n = 6 this is the binding constraint on everything below (Hastie, Tibshirani & Friedman, ESL 2nd ed., §7.10–7.12; Efron & Tibshirani, An Introduction to the Bootstrap, 1993, ch. 17).
Forecast standard deviation. FSD = sd{ln(ŷᵢ/yᵢ)}, computed on the held-out pairs. The log ratio is scale-invariant, so the statistic is comparable across price points and across vendors without normalisation. Under an approximately log-normal error, exp(±FSD) bounds a one-standard-deviation multiplicative interval; the log-normality is an approximation and is not tested at n = 6. This is the AVM industry's disclosure convention, which is what makes the comparison against the vendors’ published figures like-for-like in definition — though not in sample, since theirs are national aggregates over millions of transactions and this is six comparables for one house.
Interval construction, and why not a normal-theory interval. The 68% and 95% figures are empirical quantiles of the absolute percentage error distribution {|ŷᵢ−yᵢ|/yᵢ}, not μ̂ ± zσ̂. A normal-theory interval would import a distributional assumption the sample cannot support and would be symmetric in levels, which a multiplicative error process is not. The trade-off is explicit: a quantile at the 95th percentile of a handful of observations is an extrapolation of the fitted distribution, not an observed exceedance rate, and is labelled as such wherever it is quoted. Coverage validity requires exchangeability between the held-out comparables and the subject — the assumption underpinning conformal prediction (Vovk, Gammerman & Shafer, Algorithmic Learning in a Random World, 2005) — which close, same-era selection supports without guaranteeing. It is worth noting that conformal width calibration was measured for this pipeline and rejected: a six-month-forward holdout violates exchangeability outright, and coverage did not improve.
RMSLE. RMSLE = [n⁻¹ Σ (ln(1+ŷᵢ) − ln(1+yᵢ))²]1/2. Squaring in log space penalises proportional rather than absolute error and bounds the leverage of a single large residual relative to RMSE. It is asymmetric in levels: under-prediction is penalised more heavily than over-prediction of the same absolute size, since |ln(1−δ)| > ln(1+δ). The 1+ offset is vestigial at these magnitudes. RMSLE is a relative statistic with no absolute interpretation — it exists to rank methods, and quoting it alone says nothing.
MdAPE and MAPE. MAPE = n⁻¹ Σ |ŷᵢ−yᵢ|/yᵢ; MdAPE is the median of the same summands. MAPE is undefined at y = 0 and structurally asymmetric: unbounded above for over-prediction, bounded by 1 for under-prediction, so it systematically favours methods that under-forecast. The median is reported alongside because it has a 50% breakdown point against the mean's 1/n. Here 19.85% against 15.08% indicates one residual pulling the mean away from the median.
Hit rates. PPEk = n⁻¹ Σ 1{|ŷᵢ−yᵢ|/yᵢ ≤ k} at k = 0.10 and 0.20 — the empirical CDF of absolute percentage error at two points. On this property: 2 of 6 within 10%, 4 of 6 within 20%. Reported as counts, not percentages: the estimator takes only 7 values at n = 6, and a Wilson score interval on 2/6 spans roughly [0.10, 0.70].
What is not claimed, stated plainly. No standard error is attached to any statistic here. The sampling distribution of a standard deviation at this sample size is severely right-skewed, and a nonparametric bootstrap resamples the same few points — it would produce an interval, and that interval would be an artefact of the resampling rather than evidence about the population. No hypothesis is tested and no p-value is reported, because there is no null here worth rejecting at this sample size. These figures are offered as descriptive statistics of this comparable set, which is what they are and the most that six observations can support. The corresponding population-level figures — MdAPE 7.77%, RMSLE 0.1605 over a blind holdout of 19,486 sales — exist and are the right basis for any claim about the method in general; they are not what this page reports, because the numbers asked for are the ones specific to this property.
What others say
None of these were used to produce our number.
They are here so you can see what else is being said about this house.
One adjustment makes the comparison fair. Every figure in this report is a
net price — what a seller keeps. The outside estimates are contract
prices. So ours is shown here as a contract price too:
$505,801, which is our net $502,747 grossed back
up using the ratio Serenada West sales actually run at.
| Source | Value | Range | vs ALEX | What it is |
|---|---|---|---|---|
| ALEXthis report | $505,801$502,747 net | $331,866–$774,559 | — | Six sales, each adjusted to this property across 4 factors and reconciled by adjustment-weighted mean. |
| HouseCanaryconfidence 94%, HIGH | $542,761 | $509,129–$576,393 | +7.3% | A national automated valuation, quoting its own forecast standard deviation of 6%. Retrieved with this valuation run. |
| Zillow Zestimate | $526,500 | $494,910–$563,355 | +4.1% | Zillow rates its own confidence here “Low”, forecast standard deviation 6%. It cannot see the interior, the condition, or the subdivision boundary that governs this comparable set. |
| Williamson CADparcel R045830 | $542,517 | — | +7.3% | 2026 appraised value: land $192,386 plus improvement $350,131; prior year $536,291. A tax assessment, not an opinion of market value — produced en masse for taxation, lagging by design, and capped by statute. |
| Cotalityformerly CoreLogic | pending | — | — | Reserved — contract renewal pending. No valuation is retrieved until the agreement is executed. This is a commercial position, not a gap in this property’s data and not a failure of the model. Nothing is estimated in its place, and this row will carry a real figure when access resumes. |
There is a floor, around ±5%. No model can reliably predict a single sale much closer, because sales themselves disagree — the same house listed twice in one month will not fetch the same price. Different buyers, different negotiation, different day. That scatter is how houses are sold, not a flaw in any model, and more data will not remove it. Treat any accuracy claim much tighter than ±5% on one property with suspicion.
Where that floor is going. MIT has suggested to us that the 5% floor has not been statistically breached by anyone, but that they believe it is possible to get under 3%. We are relying on their guidance to set a new global standard. That is a target, not a claim about this report: nothing here is measured against 3%, and this page reports what it actually achieved on this property.
The vendor models are proprietary. Zillow, HouseCanary and Cotality — the company that acquired CoreLogic — publish neither their methods, their parameters nor their data, and their accuracy figures are self-scored on samples they choose. Nobody outside those companies can reproduce a number, test an interval, or prove any of it wrong — which is the minimum before a figure counts as a defensible statistic. Good commercial products; not evidence in the sense this report uses the word.
This does not make the vendor numbers useless. We use them as benchmarks. We do not include them in our valuation model.
Why an undisclosed model is not a defensible statistic. A figure is defensible when someone else can reproduce it: the estimator is specified, the data are identified, and the validation protocol is stated. Both vendors publish outputs and accuracy summaries; neither publishes the estimator, the trained parameters, or the data. Their accuracy figures are therefore self-reported and self-scored on samples they select. This is normal commercial practice and not a criticism of their competence — it is a statement about what an outside reader can check, which is: the claim, but not the calculation.
Zillow, in its own words. The Zestimate is described as combining “public records, MLS data and user-submitted home details into Zillow’s proprietary home valuation model,” offered as “a transparent, free starting point,” and explicitly “not an appraisal and can’t be used in place of an appraisal.” Published accuracy: nationwide median error 1.79% on-market and 7.20% off-market — meaning half of on-market Zestimates fall within 1.79% of the eventual sale price. Zillow states directly that on-market estimates are more accurate “because more up-to-date information is available, including listing details and recent market activity.” No specification, feature list, or training data is published.1
HouseCanary, from its published technical brief. Considerably more is disclosed. The algorithm runs in three stages: (1) query and clean data; (2) build localised price indices; (3) train machine-learning models on time-adjusted historical prices. Models are fitted at census-tract level, borrowing from neighbouring tracts where a tract is too thin to model alone, with neighbours used for training but excluded from accuracy testing. Two index models are built — median price and median price per square foot — at census-block level, and all historical sales and list prices are carried to current values through them. The response variable is each price expressed as a percent deviation from its block’s current median for that property type; several ML models per tract are then fitted to explain that deviation and combined into one estimate. Inputs include 3,100+ county assessors, 2,700+ county recorders over 20 years, MLS characteristics, listed prices and contract prices, mortgage balances and distress measures. In non-disclosure states such as Texas, MLS contract prices substitute for recorded sale prices where an arm’s-length sale can be jointly verified — directly relevant here.2
HouseCanary’s FSD, and an important caveat they state themselves. FSD is trained on the census-tract empirical error distribution and depends both on that spread and on how much the component models disagree for the individual property. The confidence score is simply 1 − FSD, which is why the 94% confidence reported for this property corresponds to an FSD of 0.06. The interval is constructed for approximately 68% coverage: P(1 ± FSD). And in their own words: “We assume symmetry but not normality... using 2*FSD to estimate a new interval is not guaranteed to yield an approximate 95% coverage probability.” Anyone doubling their FSD to get a 95% band is doing something the vendor explicitly warns against.
What HouseCanary validates, and how. Monthly internal testing on a rolling six-month window, plus quarterly blind third-party testing. As of July 2019: national MdAPE 2.8% over 1,994,203 transactions internally, and 2.9% on the third-party blind sample for Q1 2019. They report hit rate, MdAPE, median signed error (as a bias check), within-5/10/20%, and the share of sales falling inside their own 68% interval — a coverage check, which is the right diagnostic and more than most vendors publish. Results are available by state and MSA on request, which is the boundary of the disclosure: the validation design is public, the validation data are not.
The comparison this page makes, and its limit. FSD is a shared convention, so this report’s 33.21% and theirs are the same quantity — the standard deviation of log(estimate ÷ outcome). But they are computed on different samples by different methods: theirs across millions of national transactions, this one across six comparables for this house. A national median error says what happens to a typical home; it does not say what happens to this particular property and its own comparables. Reading a vendor figure as a promise about this house would be a category error, and reading ours as vendor-grade inaccuracy would be the same error in reverse.
What this report exposes by contrast. The estimator is stated, the comparables are named with MLS numbers, every adjustment is itemised with its reason, the index method is a published 1963 regression with its exclusions listed, and the held-out test results are printed so every summary statistic can be recomputed by hand. Whether the answer is right is a separate question; whether it can be checked is not.
1 zillow.com/z/zestimate — quotations from Zillow’s published “What is a Zestimate?” page. · 2 HouseCanary Valuation Technical Brief, © 2019 HouseCanary Inc., read in full. Figures are as of that publication and may have moved since. · 3 cotality.com (formerly CoreLogic) — no valuation estimate was obtained; the contract is pending renewal.