Comparable-sales valuation
Serenada West Sec 4 · Georgetown, Texas 78628 · Williamson County · MLS 3434023
The property
Land is priced at this parcel’s own Williamson CAD rate — $124,254 across 23,087 square feet, or $5.38 per square foot — not a county average and not a third-party record. The county roll reconciles for every year on file: land plus improvement equals total.
Limits
This is a comparable-sales opinion of value, not a certified appraisal, and it is not a mortgage or lending document. It is produced from the ABOR feed, Williamson CAD records and ALEX's own adjustment engine, run automatically with no comparable selected by hand. Three comparables is the minimum from which a range can honestly be quoted — the set here is tight and consistent, but it is three sales, and a market this thin deserves that said plainly rather than buried.
Competition
Two properties. Reference only — they are not used
in the valuation and carry no weight in it. Every figure in this report comes from
closed sales — properties where money actually changed hands. The two listings below have
not sold. An asking price is a seller’s opinion of value, not a measurement of it, and no
asking price enters this calculation at any point.
They are shown because they are useful context for a decision, not for a valuation: a buyer
standing in this house can see both of these the same afternoon, and they are the alternatives
this property competes against. Both appear on the
map above as blue diamonds on dashed lines. (Both also happen to be about 22% larger, which would put them
outside the comparable set’s size tolerance in any case — but the reason they are
excluded is simpler than that. They have not sold.)
| Listing | Distance | Asking | Living area | $/sf | Built | Days on market |
|---|---|---|---|---|---|---|
| 4319 Miramar Drthis property | — | $540,000 | 2,043 sf | $264.32 | 1993 | 40 |
| 300 W Sequoia Spureast-north-east | 0.47 mi | $509,999 | 2,496 sf | $204.33 | 1979 | 42 |
| 205 E Esparada Dreast | 1.08 mi | $675,000 | 2,500 sf | $270.00 | 1974 | 82 |
300 W Sequoia Spur
$509,999 · 2,496 sf · built 1979 · 0.47 mi ENE · 42 days on market
The direct threat. It asks $30,001 less than this property while offering 453 more square feet — $204 per foot against $264. A buyer weighing space against age and condition has an easy comparison to make, and it favours this listing on both price and size. Forty-two days without a sale says the market has not agreed yet, but it anchors what a buyer thinks $500,000 should buy in Serenada.
205 E Esparada Dr
$675,000 · 2,500 sf · built 1974 · 1.08 mi E · 82 days on market
The ceiling, and a cautionary one. At $270 per foot it asks more per foot than this property does, on a house nineteen years older, and it has sat 82 days without selling. It is useful mainly as evidence of what the market is declining to pay: the top of this street is being tested and is not clearing.
If you are counting against the MLS, you will find three, not two. The source data carries a third active record at 300 W Sequoia Spur — a $2,700 per month lease, listed at the same address as the sale above it. A rental is not competition for a purchase, so it is excluded here and from the map. It is named rather than quietly dropped, so the count can be reconciled: three active records, one of them a lease, two properties genuinely for sale.
To be explicit, because it matters: nothing on this page was adjusted, weighted or influenced by either of these listings. They appear in this section and on the map, and nowhere else in the calculation. If both were withdrawn tomorrow the valuation would be unchanged to the dollar. What they tell you is what this property is being shopped against — which is a question about strategy, not about value.
What these numbers are
$464,529 is what we believe the seller keeps at closing — not what the contract would read. It is the value after the seller has paid the buyer’s closing costs and made the repairs the buyer asked for. Every comparable sale below is stated the same way, so like is compared with like.
Both deductions, from this report’s own evidence
| 4301 Verde Vis — a repair credit | |
| Closed price | $545,000 |
| Buyer closing costs paid by the seller | none |
| Repairs at the buyer’s request | −$800 |
| Net sale price used here | $544,200 |
| 303 Sequoia Spur — a closing-cost concession | |
| Closed price | $550,000 |
| Buyer closing costs paid by the seller | −$300 |
| Repairs at the buyer’s request | none |
| Net sale price used here | $549,700 |
The third comparable, 609 Esparada Dr, had neither, so its net and closed price are both $482,500. Concessions on this set are small — but they are not always. Two houses can close at the same contract price and not be the same sale: if one seller paid $12,000 toward the buyer’s costs and the other paid nothing, the first achieved $12,000 less. Working in net sale price removes that difference before anything is compared, which is why these comparables are directly comparable to one another and to this property.
Asking price
Listed 7 August 2026 at $559,000, reduced to $549,900, then to $540,000 on 20 August. The date of the first reduction is not carried in the MLS feed.
| List price | Reduction, $ | Reduction, % | |
|---|---|---|---|
| Original, 7 August 2026 | $559,000 | — | — |
| First reduction, date not in feed | $549,900 | −$9,100 | −1.63% |
| Current, 20 August 2026 | $540,000 | −$9,900 | −1.80% |
| Total since listing | −$19,000 | −3.40% | |
| Per square foot | $264.32 | −$9.30 | −3.40% |
Where everything is
Everything is within about a mile, and every marker is clickable — a red circle jumps to that sale in the table below, a blue diamond to the for-sale property in its own section. Solid red lines join the three comparables to the subject; dashed blue lines join the two properties currently for sale, which are shown for reference and are not used in the valuation.
4319 Miramar Dr
4301 Verde Vis
609 Esparada Dr
303 Sequoia Spur
300 W Sequoia Spur · for sale
205 E Esparada Dr · for sale
Comparable sales — used in the valuation. 609 Esparada Dr 0.11 mi south-south-east, 4301 Verde Vis 0.29 mi east, and 303 Sequoia Spur 0.42 mi east-north-east. All three sold, all three inside Serenada West, all three within half a mile. They are joined to the subject by solid red lines, and each one carries the adjustments set out later in this report.
Properties for sale — reference only, not used in the valuation. 300 W Sequoia Spur 0.47 mi east-north-east, and 205 E Esparada Dr 1.08 mi east. Neither has sold, so neither is evidence of value; they are joined to the subject by dashed blue lines to mark that difference. 205 E Esparada is more than twice as far from this house as any comparable and sits outside the subdivision the valuation draws on. Both are shown because a buyer viewing this property will also be viewing them: see Other Properties for Sale.
Not Market Averages
Nothing here is a market average. Every figure is computed from this property’s own three comparable sales — how far the method missed when each was hidden and re-predicted, how much the three disagree with one another, how far they had to be moved to match this house, and how old they are. Open either panel below: the first states each measure in plain terms, the second gives the formal treatment.
1. Leave-one-out back-test (mean absolute percentage error). For i = 1..n the full pipeline — comparable adjustment across all 66 factors, then adjustment-weighted reconciliation — is refit on D \ {(xᵢ, yᵢ)} and evaluated at xᵢ. The reported statistic is n⁻¹ Σ |ŷ₋ᵢ − yᵢ| / yᵢ. LOOCV is approximately unbiased for expected out-of-sample risk but has high variance: the n training sets share n−2 observations, so fold estimates are strongly positively correlated and the apparent precision is optimistic (Hastie, Tibshirani & Friedman, ESL 2nd ed., §7.10). At n = 3 this is the binding caveat and the reason the figure is presented as descriptive of this comparable set rather than as a population claim.
2. Adjusted spread — a dispersion measure, not an error measure. (max vᵢ − min vᵢ) / V̂ over the adjusted comparable values. It quantifies agreement among estimates, which is not the same quantity as distance from the truth and can move in the opposite direction. Three comparables drawn from one subdivision, adjusted by one rate table, share method and data; their errors are positively correlated, so a tight spread partly reflects common specification rather than convergence on the right answer. Formally, for estimates with pairwise error correlation ρ, the variance of their mean is (σ²/n)[1 + (n−1)ρ] — under strong positive ρ the effective sample size is far below n, and observed spread understates true uncertainty. Low spread with a shared bias is the classic route to a confident wrong answer, which is why this page reports spread and held-out error separately and never substitutes one for the other.
3. Mean absolute adjustment — a proxy for accumulated model error. n⁻¹ Σ gᵢ/Pᵢ, gross adjustment over net sale price. Every adjustment is itself an estimate carrying its own error, so total adjustment magnitude bounds how much of a final figure is model rather than observed transaction. Under the approximation that adjustment errors are independent with variance proportional to magnitude, the error contributed grows with √Σg² — which is the premise the reconciliation weighting rests on, and the reason gross rather than net is the right summary. It is an ordering statistic, not a calibrated error: it ranks comparables by how much work was done to them without asserting how wrong that work was.
4. Mean comparable age — exposure to temporal drift. Mean days between each comparable’s close date and the valuation date. Time adjustment removes the expected market movement, so residual risk is the variance of the index estimate over the carry interval, not the movement itself. Under the Case–Shiller error decomposition that variance grows roughly linearly in the interval, Var ≈ σ²ᵘΔt + 2σ²ᵤ, so age is a proxy for how much index uncertainty each comparable imports. A 145-day mean is modest; the relevant exposure is not the elapsed time but the month-to-month volatility of the county index across it, which the market-timing section quantifies directly.
What the four are jointly evidence of. Taken together they describe the provenance of the estimate, not its accuracy: how much of the final figure is observed transaction and how much is model (mean absolute adjustment), whether the independent indications agree (adjusted spread), how much temporal extrapolation was required (comparable age), and how the method performed when it could not see the answer (the back-test). Only the last is an error measure. The first three can all look excellent on a valuation that is wrong.
Why they cannot simply be combined. A reader’s instinct is to average them into a single quality number. They are not independent, and the dependence runs the wrong way. Comparables selected for similarity have small adjustments by construction — the cascade chose them for it — so a low mean adjustment partly reflects the selection rule rather than the market. Tight agreement among three comparables adjusted by one rate table reflects the common rate table. Formally, for estimates with pairwise error correlation ρ, the variance of their mean is (σ²/n)[1 + (n−1)ρ]: at ρ near 1 the effective sample size approaches 1 regardless of n, and observed spread understates true uncertainty by an unknown factor. Nothing here estimates ρ — three observations cannot.
The failure mode this section exists to expose. Comparables that are genuinely alike, adjusted by a shared and slightly wrong rate table, will produce a tight spread, small adjustments, and a low back-test error — while every one of them is biased in the same direction. Correlated error is invisible to all four measures simultaneously. That is not a hypothetical on this property: two comparables sit on double the land, both take a large adjustment from the same land rate, and if that rate is wrong both move together. It is the reason the land section states its rate and its source rather than burying them.
What none of these four is. None is a confidence interval, and none should be read as one. Coverage statements live in Cross Validation, where they are derived from the held-out error distribution rather than from agreement among comparables — agreement among correlated estimates being precisely the thing that can be high while the answer is wrong.
The land
Two of the three comparables sit on roughly an acre; this property sits on half of one. That difference is priced at a rate taken from the target’s own parcel — what Williamson County says its land is worth, divided by how much land the county says it has. Nothing here is estimated or taken from a listing portal.
| Parcel | CAD acres | Lot size | CAD land value | County $/sf |
|---|---|---|---|---|
| 4319 Miramar Drthe subject, R045978 | 0.530 | 23,087 sf | $124,254 | $5.38 |
| 4301 Verde Vis | 1.020 | 44,431 sf | $192,386 | $4.33 |
| 303 Sequoia Spur | 1.010 | 43,996 sf | $191,196 | $4.35 |
| 609 Esparada Dr | 0.510 | 22,216 sf | $120,938 | $5.44 |
The county’s acreage and the MLS lot size agree exactly here — 0.530 acres is 23,086.8 square feet, and the listing says 23,086.8 — so the two independent records of this parcel’s size do not conflict.
| Comparable | Lot size | Difference | Land adjustment | Share of net |
|---|---|---|---|---|
| 4301 Verde Vis | 44,431 sf | +21,344 sf | −$114,876 | 21.1% |
| 303 Sequoia Spur | 43,996 sf | +20,909 sf | −$112,533 | 20.5% |
| 609 Esparada Dr | 22,216 sf | −871 sf | +$4,689 | 1.0% |
Market timing
A sale that closed in January is a snapshot in time. To use it as evidence about today, it has to be moved forward by however much the market moved in between. That is what the market timing line on each comparable does.
| Comparable | Closed | Index then | Index now | Factor | Adjustment |
|---|---|---|---|---|---|
| 609 Esparada Dr | 2026-01 | 166.975 | 167.172 | 1.0012 | +$568 |
| 4301 Verde Vis | 2026-05 | 166.970 | 167.172 | 1.0012 | +$656 |
| 303 Sequoia Spur | 2026-06 | 171.992 | 167.172 | 0.9720 | −$15,406 |
Adjustment = net sale price × (factor − 1). All three come from the Williamson County index, built from 19,824 repeat-sale pairs.
What the index numbers mean. They are not dollars and not percentages — they are a scale, and every scale needs a zero. Here the zero is January 2016, set to 100. That month is the base by construction: the regression fixes its value and measures every later month against it. So a reading of 167.172 for September 2026 means Williamson County house prices are 67.2% above where they stood in January 2016.
Specification. Bailey, Muth & Nourse (1963), A Regression Method for Real Estate Price Index Construction, JASA 58(304), 933–942. For a dwelling transacting at periods s and t with s < t, the model is ln(Pᵗ/Pₛ) = βᵗ − βₛ + εₛᵗ. Let X be the N × T design matrix with xᵢₛ = −1, xᵢᵗ = +1 and zero elsewhere. Then β̂ = (X′X)⁻¹X′y with yᵢ = ln(Pᵗ/Pₛ). The base period is normalised to β₀ = 0 and its column deleted; retaining it makes X′X singular, since the indicators sum to zero by construction. The index is Iₜ = 100·exp(β̂ₜ).
Identification. The estimator is consistent for the market factor under the assumption that E[ε | X] = 0 — that idiosyncratic price changes are mean-independent of when a property happened to transact. Dwelling fixed effects difference out exactly because the same unit appears with opposite signs on both sides, which is what removes omitted-variable bias from unobserved quality and is the method’s central advantage over a hedonic index.
Error structure. Case & Shiller (1987, New England Economic Review; 1989, AER 79(1), 125–137) decompose εₛᵗ = (Hᵗ − Hₛ) + (uᵗ − uₛ), where H is a Gaussian random walk in the property’s own value and u is i.i.d. transaction noise. The implied variance Var(ε) = σ²ᵘ(t−s) + 2σ²ᵤ grows linearly in the holding interval, motivating their three-stage weighted least squares: OLS, regress squared residuals on the interval, then GLS with weights 1/√(â + b̂(t−s)). This index does not apply that weighting. It instead excludes holds under 180 days and annualised changes beyond ±60% outright — a trimming rather than a down-weighting strategy, which is more conservative, discards information the WLS would retain, and is stated rather than left implicit.
Shrinkage. County-month estimates are shrunk toward the metropolitan estimate in log space, ln Iᶜₜ = w ln Iᶜₜᶜ + (1−w) ln Iᵐₜ with w = n/(n+k), k = 20. This is the James–Stein / empirical-Bayes posterior mean under a normal-normal hierarchy where k is the ratio of within-month sampling variance to between-county prior variance; here it is fixed rather than estimated. The effect is that thin months borrow strength in proportion to how little evidence they carry of their own — September 2026, with 38 pairs, sits at w = 0.655.
Known limitations. (i) Sample selection. The index is estimated only on dwellings that transacted at least twice, which is not a random sample of the stock; properties that turn over frequently differ systematically, and no Heckman-type correction is applied. (ii) Revision. Estimates for recent periods move as later transactions arrive, so the most recent months are the least stable — a property shared with all repeat-sales indices including S&P CoreLogic Case–Shiller. (iii) Renovation. The dwelling-constant assumption fails where a property was materially improved between sales; the ±60% annualised filter removes the extreme cases and cannot remove the moderate ones. (iv) Aggregation. A single county index assumes a common market factor across submarkets within the county.
Empirical note on smoothing. A trailing three-month average of the log index was tested against the raw monthly series by ten-fold cross-validation on 20,631 Williamson pairs, scoring out-of-sample RMSE of predicted log price change. It was worse in 10 of 10 folds (mean difference +0.00171, 95% CI [+0.00139, +0.00202], t = +10.70); a five-month window was worse again (t = +22.77). Monotonicity in the smoothing window is the signature of signal removal rather than noise suppression, consistent with the shrinkage above already performing the variance reduction a second pass would duplicate.
Cross Validation
How this was tested, in one sentence: we covered up one of the three sales, worked out what this method would have said that house was worth using only the other two, then uncovered it and measured how close we got — and did that three times, once for each sale.
Three tests is the minimum from which a spread can be quoted. These figures are real and specific to this house, but thin: one unusual comparable moves them materially. Treat them as the shape of the uncertainty, not a precise measurement of it.
| Held out | Actual net | Predicted from the other two | Error |
|---|---|---|---|
| 4301 Verde Vis | $544,200 | $559,630 | +2.84% |
| 303 Sequoia Spur | $549,700 | $586,589 | +6.71% |
| 609 Esparada Dr | $482,500 | $415,669 | −13.85% |
Part one — what a valuation should be measured against
This section is reserved. The standard is being supplied separately and will be placed here in full. It is deliberately empty rather than filled with a plausible-looking threshold: a benchmark that this report happened to pass would be indistinguishable, to a reader, from one it had actually been measured against. Until the paper is here, no accuracy standard is being claimed or implied anywhere in this report.
What is known today, pending that standard. Three sets of numbers exist and none of them is a standard — they are measurements, and they are not measuring the same thing:
1. This property. Median error 6.71%, FSD 11.44%, RMSLE 0.0953, 68% interval ±9.28%, from three held-out comparables. Specific to this house, and thin: three observations describe this comparable set accurately and generalise weakly.
2. This model, across its blind holdout. MdAPE 7.77%, MAPE 11.22%, RMSLE 0.1605, 95% coverage 89.87% over 19,486 sales the model never saw during training. By price band on single-family homes: $250–400k 6.09%, $400–600k 7.20%, $600k–1M 8.35%, under $250k 9.92%, over $1M 13.10%. This is the right basis for any general claim about the method, and it is deliberately not what the tiles above report — those answer “how accurate is this valuation”, which is a different question.
3. What the vendors publish for themselves. HouseCanary: national MdAPE 2.8% over 1,994,203 transactions internally, 2.9% on a blind third-party test. Zillow: median error 1.79% on-market and 7.20% off-market. Two cautions before either is treated as a bar to clear. Both are national aggregates dominated by dense, homogeneous tract housing where an automated model does best, and neither is conditioned on the kind of property this is. And the on-market figure is not independent of the asking price — both vendors consume list price, which is the whole point made in the Benchmarks section.
What a standard would need to specify to be usable here, and what the forthcoming paper will presumably settle: which statistic is the criterion (MdAPE, FSD, coverage, or a hit rate); whether it is a point threshold or a distribution; whether it applies per property or per portfolio; how it is conditioned on price band, property type and evidence depth; and whether an interval must be calibrated — that is, whether a stated 68% interval must actually contain the outcome 68% of the time, which is a materially stricter requirement than a median error threshold and the one most AVM disclosures avoid.
Where this report already exceeds common practice, whatever the standard turns out to be: the comparables are named, every adjustment is itemised with its reason, the estimator and its weights are stated, and the three held-out test results are printed so every summary statistic can be recomputed by hand. A standard can be applied to a number; it can only be verified against a number whose derivation is visible.
Part two — how this valuation was measured
Why these figures are worse than an earlier version of this report. Every figure above is computed from the three held-out rows and can be reconciled by hand. They are materially worse than earlier versions because the land difference is now carried in full. When the two acre-sized comparables were adjusted by a limited amount the median error was 2.65% and the FSD 3.53%; carrying the full $114,876 and $112,533 they are 6.71% and 11.44%. Both things are true at once: the full figures are what the method produces, and the smaller ones predicted these three sales more closely. The honest reading is that two comparables on double the land are hard to adjust reliably in either direction.
Why hold anything out at all. Scoring a method on the same sales it was given measures how well it can interpolate data it has already seen, which is not the quantity anyone cares about. The quantity of interest is expected loss on an observation drawn from the same population and not used in fitting: Err = E(X,Y)[L(Y, f̂(X))]. Resubstitution error is downward-biased for it by construction. The land difference is where this bites on this property. Two of the three comparables sit on about double its land, so each time one was held out the method had to price roughly half an acre with no recent sale of its own to price it from. Those are the misses in the table above.
Design: leave-one-out cross-validation. For i = 1,…,n the entire pipeline — all 66 adjustment factors, the repeat-sales time carry, and the inverse-gross-adjustment reconciliation — is refit on D−i = D \ {(xᵢ,yᵢ)} and evaluated at xᵢ, giving CV = n⁻¹ Σᵢ L(yᵢ, f̂−i(xᵢ)). Critically, the held-out comparable is excluded from the reconciliation as well as from the adjustment grid; leaving it in the weighting while removing it from the grid would leak the answer through the weights.
Properties, and the one that governs here. LOOCV is approximately unbiased for Err — it trains on n−1 of n points, so the pessimistic bias of k-fold at small k is largely absent. It pays for this in variance: the n training sets are nearly identical, sharing n−2 observations, so the fold-level estimates are strongly positively correlated and Var(CV) does not contract at the 1/n rate an independent average would. At n = 3 this is the binding constraint on everything below (Hastie, Tibshirani & Friedman, ESL 2nd ed., §7.10–7.12; Efron & Tibshirani, An Introduction to the Bootstrap, 1993, ch. 17).
Forecast standard deviation. FSD = sd{ln(ŷᵢ/yᵢ)}, computed on the held-out pairs. The log ratio is scale-invariant, so the statistic is comparable across price points and across vendors without normalisation. Under an approximately log-normal error, exp(±FSD) bounds a one-standard-deviation multiplicative interval; the log-normality is an approximation and is not tested at n = 3. This is the AVM industry's disclosure convention, which is what makes the comparison against HouseCanary's 3% and Zillow's 5% like-for-like in definition — though not in sample, since theirs are national aggregates over millions of transactions and this is three comparables for one house.
Interval construction, and why not a normal-theory interval. The 68% and 95% figures are empirical quantiles of the absolute percentage error distribution {|ŷᵢ−yᵢ|/yᵢ}, not μ̂ ± zσ̂. A normal-theory interval would import a distributional assumption the sample cannot support and would be symmetric in levels, which a multiplicative error process is not. The trade-off is explicit: a quantile at the 95th percentile of three observations is an extrapolation of the fitted distribution, not an observed exceedance rate, and is labelled as such wherever it is quoted. Coverage validity requires exchangeability between the held-out comparables and the subject — the assumption underpinning conformal prediction (Vovk, Gammerman & Shafer, Algorithmic Learning in a Random World, 2005) — which same-subdivision, same-era selection supports without guaranteeing. It is worth noting that conformal width calibration was measured for this pipeline and rejected: a six-month-forward holdout violates exchangeability outright, and coverage did not improve.
RMSLE. RMSLE = [n⁻¹ Σ (ln(1+ŷᵢ) − ln(1+yᵢ))²]1/2. Squaring in log space penalises proportional rather than absolute error and bounds the leverage of a single large residual relative to RMSE. It is asymmetric in levels: under-prediction is penalised more heavily than over-prediction of the same absolute size, since |ln(1−δ)| > ln(1+δ). The 1+ offset is vestigial at these magnitudes. RMSLE is a relative statistic with no absolute interpretation — it exists to rank methods, and quoting it alone says nothing.
MdAPE and MAPE. MAPE = n⁻¹ Σ |ŷᵢ−yᵢ|/yᵢ; MdAPE is the median of the same summands. MAPE is undefined at y = 0 and structurally asymmetric: unbounded above for over-prediction, bounded by 1 for under-prediction, so it systematically favours methods that under-forecast. The median is reported alongside because it has a 50% breakdown point against the mean's 1/n. Here 7.80% against 6.71% indicates no single dominating residual.
Hit rates. PPEk = n⁻¹ Σ 1{|ŷᵢ−yᵢ|/yᵢ ≤ k} at k = 0.10 and 0.20 — the empirical CDF of absolute percentage error at two points. On this property: 2 of 3 within 10%, 3 of 3 within 20%. Reported as counts, not percentages: the estimator takes only four values at n = 3, and a Wilson score interval on 2/3 spans roughly [0.21, 0.94], which is most of the parameter space.
What is not claimed, stated plainly. No standard error is attached to any statistic here. The sampling distribution of a standard deviation at n = 3 is severely right-skewed, and a nonparametric bootstrap resamples the same three points — it would produce an interval, and that interval would be an artefact of the resampling rather than evidence about the population. No hypothesis is tested and no p-value is reported, because there is no null here worth rejecting at this sample size. These figures are offered as descriptive statistics of this comparable set, which is what they are and the most that three observations can support. The corresponding population-level figures — MdAPE 7.77%, RMSLE 0.1605 and 95% coverage of 89.87% over a blind holdout of 19,486 sales — exist and are the right basis for any claim about the method in general; they are not what this page reports, because George asked for the numbers specific to this property and those are these.
What others say
None of these were used to produce our number.
They are here so you can see what else is being said about this house.
One adjustment makes the comparison fair. Every figure in this report is a
net price — what a seller keeps. The three outside estimates are contract
prices. So ours is shown here as a contract price too:
$465,717, which is our net $464,529 grossed back
up using the ratio Serenada sales actually run at.
| Source | Value | Range | vs ALEX | What it is |
|---|---|---|---|---|
| ALEXthis report | $465,717$464,529 net | $411,170–$501,746 | — | Three Serenada West sales, each adjusted to this property across 66 factors and reconciled by adjustment-weighted mean. |
| HouseCanaryconfidence 97%, HIGH | $552,812 | $534,631–$570,993 | +18.7% | A national automated valuation, quoting its own forecast standard deviation of 3%. Retrieved with the valuation run of 15 September 2026. |
| Zillow Zestimatezpid 29544057 | $529,700 | $503,215–$556,185 | +13.7% | Zillow rates its own confidence here “Good”, forecast standard deviation 5%. It cannot see the interior, the condition, or the subdivision boundary that governs this comparable set. |
| Williamson CADparcel R045978 | $441,435 | — | −5.2% | 2026 appraised value: land $124,254 plus improvement $317,181; prior year $429,076. A tax assessment, not an opinion of market value — produced en masse for taxation, lagging by design, and capped by statute. |
| Cotalityformerly CoreLogic | pending | — | — | Reserved — contract renewal pending. No valuation is retrieved until the agreement is executed. This is a commercial position, not a gap in this property’s data and not a failure of the model. Nothing is estimated in its place, and this row will carry a real figure when access resumes. |
There is a floor, around ±5%. No model can reliably predict a single sale much closer, because sales themselves disagree — the same house listed twice in one month will not fetch the same price. Different buyers, different negotiation, different day. That scatter is how houses are sold, not a flaw in any model, and more data will not remove it. Treat any accuracy claim much tighter than ±5% on one property with suspicion.
Where that floor is going. MIT has suggested to us that the 5% floor has not been statistically breached by anyone, but that they believe it is possible to get under 3%. We are relying on their guidance to set a new global standard. That is a target, not a claim about this report: nothing here is measured against 3%, and this page reports what it actually achieved on this property.
The vendor models are proprietary. Zillow, HouseCanary and Cotality — the company that acquired CoreLogic — publish neither their methods, their parameters nor their data, and their accuracy figures are self-scored on samples they choose. Nobody outside those companies can reproduce a number, test an interval, or prove any of it wrong — which is the minimum before a figure counts as a defensible statistic. Good commercial products; not evidence in the sense this report uses the word.
Now the thing worth noticing. Both estimates land within 2.4% of the $540,000 asking price — HouseCanary 2.4% above, Zillow 1.9% below. This report is 13.8% below. The reason is simple: they use the asking price; we do not. HouseCanary’s own brief lists MLS “listed prices and contract prices” as inputs and says a new listing triggers an immediate re-valuation. Zillow publishes two accuracy figures for the same reason: 1.79% median error on-market against 7.20% off.
Set that 1.79% against the floor. It is about three times tighter than transaction noise alone permits — the signature of a model that already knows the asking price, not one that has out-predicted the market. So when an automated estimate hugs an asking price, much of what you are seeing is that asking price reflected back. To HouseCanary, Zillow and Cotality, a list price is real information about what a seller and their listing agent want. It is not independent evidence of value, which is what this report was asked for. Ours uses only what three neighbours actually sold for.
This does not make the vendor numbers useless. We use them as benchmarks. We do not include them in our valuation model.
Why an undisclosed model is not a defensible statistic. A figure is defensible when someone else can reproduce it: the estimator is specified, the data are identified, and the validation protocol is stated. Both vendors publish outputs and accuracy summaries; neither publishes the estimator, the trained parameters, or the data. Their accuracy figures are therefore self-reported and self-scored on samples they select. This is normal commercial practice and not a criticism of their competence — it is a statement about what an outside reader can check, which is: the claim, but not the calculation.
Zillow, in its own words. The Zestimate is described as combining “public records, MLS data and user-submitted home details into Zillow’s proprietary home valuation model,” offered as “a transparent, free starting point,” and explicitly “not an appraisal and can’t be used in place of an appraisal.” Published accuracy: nationwide median error 1.79% on-market and 7.20% off-market — meaning half of on-market Zestimates fall within 1.79% of the eventual sale price. Zillow states directly that on-market estimates are more accurate “because more up-to-date information is available, including listing details and recent market activity.” No specification, feature list, or training data is published.1
HouseCanary, from its published technical brief. Considerably more is disclosed. The algorithm runs in three stages: (1) query and clean data; (2) build localised price indices; (3) train machine-learning models on time-adjusted historical prices. Models are fitted at census-tract level, borrowing from neighbouring tracts where a tract is too thin to model alone, with neighbours used for training but excluded from accuracy testing. Two index models are built — median price and median price per square foot — at census-block level, and all historical sales and list prices are carried to current values through them. The response variable is each price expressed as a percent deviation from its block’s current median for that property type; several ML models per tract are then fitted to explain that deviation and combined into one estimate. Inputs include 3,100+ county assessors, 2,700+ county recorders over 20 years, MLS characteristics, listed prices and contract prices, mortgage balances and distress measures. In non-disclosure states such as Texas, MLS contract prices substitute for recorded sale prices where an arm’s-length sale can be jointly verified — directly relevant here.2
HouseCanary’s FSD, and an important caveat they state themselves. FSD is trained on the census-tract empirical error distribution and depends both on that spread and on how much the component models disagree for the individual property. The confidence score is simply 1 − FSD, which is why the 97% confidence reported for this property corresponds to an FSD of 0.03. The interval is constructed for approximately 68% coverage: P(1 ± FSD). And in their own words: “We assume symmetry but not normality... using 2*FSD to estimate a new interval is not guaranteed to yield an approximate 95% coverage probability.” Anyone doubling their FSD to get a 95% band is doing something the vendor explicitly warns against.
What HouseCanary validates, and how. Monthly internal testing on a rolling six-month window, plus quarterly blind third-party testing. As of July 2019: national MdAPE 2.8% over 1,994,203 transactions internally, and 2.9% on the third-party blind sample for Q1 2019. They report hit rate, MdAPE, median signed error (as a bias check), within-5/10/20%, and the share of sales falling inside their own 68% interval — a coverage check, which is the right diagnostic and more than most vendors publish. Results are available by state and MSA on request, which is the boundary of the disclosure: the validation design is public, the validation data are not.
The comparison this page makes, and its limit. FSD is a shared convention, so this report’s 11.44% and their 3% and 5% are the same quantity — the standard deviation of log(estimate ÷ outcome). But they are computed on different samples by different methods: theirs across millions of national transactions, this one across three comparables for this house. A national median error says what happens to a typical home; it does not say what happens to a half-acre property whose nearest comparables sit on twice the land. Reading their 3% as a promise about this house would be a category error, and reading our 11.44% as vendor-grade inaccuracy would be the same error in reverse.
What this report exposes by contrast. The estimator is stated, the comparables are named with MLS numbers, every adjustment is itemised with its reason, the index method is a published 1963 regression with its exclusions listed, and the three held-out test results are printed so every summary statistic can be recomputed by hand. Whether the answer is right is a separate question; whether it can be checked is not.
1 zillow.com/z/zestimate — quotations from Zillow’s published “What is a Zestimate?” page. · 2 HouseCanary Valuation Technical Brief, © 2019 HouseCanary Inc., read in full. Figures are as of that publication and may have moved since. · 3 cotality.com (formerly CoreLogic) — no valuation estimate was obtained; the contract is pending renewal.