The request sheet. This is the form an agent fills in to ask for the valuation below — shown here so the whole package reads as one document. Wording is editable like the rest of the page; the Send button is off on this copy, and the live sheet is at /valuations/request.

ALEX · Valuation

Request a valuation

What the engine needs before it can run: who it is for, in what capacity, and which property. Everything else the report works out for itself. A request is queued and a report is produced against it; nothing here commits a number.

Who it is for

Client capacity

The capacity decides what the report shows: a buyer sees offers toward the asking price, a seller sees pricing postures, and an off-market valuation carries no listing sections.

Which property

Pick a match and the city, state and ZIP come with it. An MLS number works here too.

What this market does

Reading the market…

Who is asking
What you already know

You have seen the house and you know the loan. These two set the starting contract price; the report measures its own defaults from the market and uses yours instead when you give them.

Requests are stored and worked in order. Nothing on this page is a valuation, and no figure here is published to a client.

Comparable-sales valuation

4319 Miramar Drive

Serenada West Sec 4 · Georgetown, Texas 78628 · Williamson County · MLS 3434023

$464,529
Comparable analysis value · net sale price
Supported range$411.2k–$501.7kLow and high of the comparables as ALEX adjusted them, before any agent change
Asking price$540,000Property has been reduced
Confidence — 68% error band0.0990This property’s own 68% band, as a decimal: the true figure should land within ±9.90% about two times in three. It is the 68th percentile of the held-out absolute errors, not the forecast standard deviation: the FSD on this property is 0.1224, and on the vendor convention of 1 − FSD, 0.8776. Evidence quality scores 60 of 100 — medium; what that score is built from is set out in Cross Validation.
Evidence3 salesAll within Serenada West

Comparable-sales valuation · 18 September 2026, 5:48 PM CDT · v14.0 · 453-tier cascade · comps agent-selected

Summary

The key figures

ALEX indicated $464,529, a net sale price, from 3 closed sales, all within Serenada West; adjusted comparables range $411,170 to $501,746. Tested on this property by leaving each comparable out: median error 6.8%, largest 15.4%; 68% interval ±9.9% ($418,528 to $510,531), 95% interval ±14.5%. Evidence quality 60 of 100 (medium). Asking $540,000 after 2 reductions from $559,000 (−3.4%), 42 days on market; the value is 14.0% below asking. Comparable homes in MLS area GTW (Georgetown) closed at a median 98.0% of asking (521 sales, 7 August 2023 to 4 August 2026). WCAD 2026 market value $441,435, +2.9% from 2025.

The property

4319 Miramar Drive

Living area2,043 sfOne story
Bedrooms / baths4 / 2Two-car garage
Year built1993Newer than most of Serenada West
Lot0.53 ac23,087 sf
Days on market42Listed 7 August 2026
WCAD 2026 market$441,435Land $124,254 + improvement $317,181
Annual tax$6,6751.5121% effective
ParcelR045978Serenada West Sec 4, Block D, Lot 29

Land is priced at this parcel’s own Williamson CAD rate — $124,254 across 23,087 square feet, or $5.38 per square foot — not a county average and not a third-party record. The county roll reconciles for every year on file: land plus improvement equals total.

Limits

What this is

Please read

This is a comparable-sales opinion of value, not a certified appraisal, and it is not a mortgage or lending document. It is produced from the ABOR feed, Williamson CAD records and ALEX's own adjustment engine, from comparables selected by the agent. Three comparables is the minimum from which a range can honestly be quoted.

Competition

Other Properties for Sale

One property. Reference only — it is not used in the valuation and carries no weight in it. Every figure in this report comes from closed sales — properties where money actually changed hands. The listing below has not sold. An asking price is a seller’s opinion of value, not a measurement of it, and no asking price enters this calculation at any point.

It is shown because it is useful context for a decision, not for a valuation: a buyer standing in this house can see it the same afternoon, and it is the alternative this property competes against. It appears on the map below as blue diamonds on dashed lines. (It also happens to be more than 20% different in size, which would put it outside the comparable set’s size tolerance in any case — but the reason it is excluded is simpler than that. It has not sold.)

ListingDistanceAskingLiving area $/sfBuiltDays on market
4319 Miramar Drthis property $540,0002,043 sf $264.32199342
205 E Esparada Dreast 1.09 mi$675,0002,500 sf $270.00197484

205 E Esparada Dr

$675,000 · 2,500 sf · built 1974 · 1.09 mi E · 84 days on market

It asks $135,000 more than this property while offering 457 more square feet — $270 per foot against $264. Built 19 years earlier; on the market 84 days without selling.

If you are counting against the MLS, you will find two, not one. The source data also carries 300 W Sequoia Spur at $2,700 — priced as a monthly rent, not a sale. A rental is not competition for a purchase, so it is excluded here and from the map, and named rather than quietly dropped so the count can be reconciled.

To be explicit, because it matters: nothing on this page was adjusted, weighted or influenced by this listing. It appears in this section and on the map, and nowhere else in the calculation. If it were withdrawn tomorrow the valuation would be unchanged to the dollar. What it tells you is what this property is being shopped against — which is a question about strategy, not about value.

What these numbers are

Every figure in this report is a net sale price

36%
Gave nothing at allOf the 521 comparable sales in this pool, this share closed with no seller contribution to the buyer's closing costs and no repair credit — the contract price and the net were the same number.
72%
Cut the asking price firstThis share had already reduced the asking price before the sale happened. This property has.

And when they did pay, what they paid

51%
Paid the buyer's closing costsThis share of the same pool put something toward the buyer's closing costs. A contribution is a percentage of the contract price, so it scales with every offer.
22%
Credited repairsThis share credited repairs. A repair credit is a fixed amount out of the inspection, so it does not scale with the price.

Same pool, same window as above: MLS area GTW, Georgetown, single family residence, asking $459,000 to $621,000, closed 7 August 2023 through 4 August 2026 — 521 sales, effective n 249.4. The two shares overlap: a sale can carry both, which is why they do not sum to the share that paid anything at all.

Pool: single family residence resales in MLS area GTW, asking between $459,000 and $621,000, closed 7 August 2023 to 4 August 2026. Recent sales count more heavily, so these are weighted shares.

Click here for full statistics — the four shares above

What is being estimated. Four proportions in one pool of closed sales: p₀ = Pr(C + R = 0), the share where the seller paid nothing toward the buyer’s closing costs and credited no repairs, and pᵣ = Pr(original list > list), the share that had reduced the asking price before the sale; and the same pool split the other way, pᶜ = Pr(C > 0), the share where the seller put something toward the buyer’s closing costs, and pᵣ = Pr(R > 0), the share that credited repairs. All four are properties of this pool, not forecasts for this property. pᶜ and pᵣ are not mutually exclusive — a sale carrying both is counted in each — so they do not sum to 1 − p₀ and should not be read as a partition.

Weighting, and why the interval is not built on 521. Sales are weighted by recency on a 6-month half-life, wᵢ = 2^(−aᵢ/h) with aᵢ the age of the sale in months, so the estimate is the weighted proportion p̂ = Σwᵢxᵢ / Σwᵢ. Weighting costs information: the effective sample size is nᴱ = (Σwᵢ)² / Σwᵢ² = 249.4 against 521 rows, and every interval below uses nᴱ. Using the row count instead would overstate precision by about 1.4 times.

Intervals. Wilson score intervals at 95%, which is the right choice here because a share near 0 or 1 makes the normal approximation cover badly and can put a bound outside [0, 1]: gave nothing 35.9% [30.2%, 42.0%]; had cut the price 72.3% [66.4%, 77.5%]; paid closing costs 50.6% [44.4%, 56.7%]; credited repairs 22.2% [17.5%, 27.8%]. The intervals assume the sales are independent. They are not entirely: sales cluster by subdivision, by listing agent and by month, so a cluster-robust interval would be wider. Read these as the narrowest defensible bounds, not the final word.

Measurement limits, stated rather than buried. A zero and a missing value are indistinguishable in the concession fields, so p₀ is an upper bound on how often a seller truly gave nothing. The reduction test compares the feed’s original list price with its current one, and a relisting can reset that original, so pᵣ is a lower bound on how often an asking price came down. Neither is corrected by a guess.

Pool and exclusions. Single Family Residence resales in MLS area GTW, asking price within ±15% of this property’s, at least 15 days on market, closed 7 August 2023 to 4 August 2026 with a 45-day settle window so pending sales cannot enter, and a 36-month cap on age. Sun City and other 55+ communities are excluded, as is new construction: both negotiate on terms a resale does not.

$464,529 is what ALEX indicates this house is worth with any seller credit stripped back out, and Diane Hart Alexander signed off $505,194 — the figure this report carries — not what the contract would read. It is the value after the seller has paid the buyer’s closing costs and made the repairs the buyer asked for. Every comparable sale below is stated the same way, so like is compared with like.

Net sale price = Closed price
Buyer closing costs paid by the seller
Repairs made by the seller at the buyer’s request
Closed price What the contract says. The headline number in the MLS, and the one most reports stop at.
Buyer closing costs paid by the seller A seller concession. Money that never reaches the seller, so counting it as sale price overstates what the house achieved.
Repairs at the buyer’s request Work or credits agreed after inspection. Also money out, also not value the next buyer is paying for.

Both deductions, from this report’s own evidence

4301 Verde Vis — a repair credit
Closed price$545,000
Buyer closing costs paid by the sellernone
Repairs at the buyer’s request−$800
Net sale price used here$544,200
303 Sequoia Spur — a closing-cost concession
Closed price$550,000
Buyer closing costs paid by the seller−$300
Repairs at the buyer’s requestnone
Net sale price used here$549,700

The other comparable, 609 Esparada Dr, had neither, so its net and closed price are both $482,500. Two houses can close at the same contract price and not be the same sale: if one seller paid $12,000 toward the buyer’s costs and the other paid nothing, the first achieved $12,000 less. Working in net sale price removes that difference before anything is compared, which is why these comparables are directly comparable to one another and to this property.

Asking price

Reduced twice since listing

Listed 7 August 2026 at $559,000, reduced to $549,900, then to $540,000 on 20 August. The date of the first reduction is not carried in the MLS feed.

$559,000 Original
7 Aug 2026
$549,900 First cut
date not in feed
$540,000 Current
20 Aug 2026
$500,000 — chart floor
 List priceReduction, $Reduction, %
Original, 7 August 2026$559,000
First reduction, date not in feed$549,900 −$9,100−1.63%
Current, 20 August 2026$540,000 −$9,900−1.80%
Total since listing −$19,000−3.40%
Per square foot$264.32−$9.30−3.40%

Where everything is

The properties

Everything is within about a mile, and every marker is clickable — a red circle jumps to that sale in the table below, a blue diamond to the for-sale property in its own section. Solid red lines join the three comparables to the subject; dashed blue lines join the one property currently for sale, which is shown for reference and is not used in the valuation.

4319 Miramar Drive — the subject Comparable sale — used in the valuation (click to jump) For sale now — blue diamond, dashed line, reference only

Comparable sales — used in the valuation. 609 Esparada Dr 0.11 mi south-south-east, 4301 Verde Vis 0.29 mi east, and 303 Sequoia Spur 0.42 mi east-north-east. All three sold, all three inside Serenada West, all three within half a mile. They are joined to the subject by solid red lines, and each one carries the adjustments set out later in this report.

Properties for sale — reference only, not used in the valuation. 205 E Esparada Dr 1.09 mi east. It has sold, so it is not evidence of value; it is joined to the subject by dashed blue lines to mark that difference. 205 E Esparada Dr is more than twice as far from this house as any comparable and sits outside the subdivision the valuation draws on. It is shown because a buyer viewing this property will also be viewing it: see Other Properties for Sale.

Schools

Schools named on the listing

The listing names Georgetown ISD, rated B- (80) by the Texas Education Agency for 2025-26, up 4 points on the year. Elementary: Raye McCoy Elementary, B- (80), top 54% of Texas elementary campuses, #7 of 11 in the district; about 1.4 miles, 3 minutes by car. Middle: Charles A Forbes Middle, C (75), top 74% of Texas middle campuses, #3 of 4 in the district; about 5.4 miles, 12 minutes by car. High School: Georgetown High, B+ (89), top 50% of Texas high school campuses, #1 of 3 in the district; about 4.8 miles, 12 minutes by car.

LevelSchoolTEA 2025-26StatewideTrendDistrict rank
DistrictGeorgetown ISD22 campuses, 22 rated B- · 80 +4
ElementaryRaye McCoy ElementaryAchievement B · Progress C · Gaps C B- · 80 Top 54% +4 #7 of 11level avg 81
MiddleCharles A Forbes MiddleAchievement C · Progress C · Gaps C C · 75 Top 74% +1 #3 of 4level avg 79
High SchoolGeorgetown HighAchievement A · Progress B · Gaps B B+ · 89 Top 50% +3 #1 of 3level avg 84

The campuses, 2024-25

SchoolStudentsEcon. disadvantagedStudents per teacherTeacher experienceFrom the property
Raye McCoy Elementary 486 23.5% 12.0 12.8 yrs 1.4 mi~3 min drive
Charles A Forbes Middle 803 46.1% 15.2 9.8 yrs 5.4 mi~12 min drive
Georgetown High 2,059 23.9% 15.5 14.5 yrs 4.8 mi~12 min drive

Outcomes, 2024-25

SchoolSTAAR at meets grade levelAttendanceChronically absent4-year graduationIn Texas college
Raye McCoy Elementary 50% 95.3% 10.0%
Charles A Forbes Middle 41% 94.7% 14.4%
Georgetown High 61% 94.6% 14.9% 98.7% 48.6%
Georgetown ISD (district) B- · 80Elementary: Raye McCoy Elementary B- · 80Middle: Charles A Forbes Middle C · 75High School: Georgetown High B+ · 89

TEA accountability score out of 100, 2025-26. The district is in red, each campus named on the listing below it.

These are the schools the listing agent named on the MLS record, matched to Texas Education Agency campus records by name within the district. This is not an attendance-zone determination: zones are set by the district and change. Verify enrollment with the district.

Click here for sources and method — the school ratings

What “top N%” is measured against. Elementary: 3,911 rated Texas campuses; Middle: 1,123 rated Texas campuses; High School: 1,352 rated Texas campuses. A percentile is the share of those campuses scoring at or above this one in 2025-26. No interval is attached: the comparison is a census of rated campuses for that year, not a sample, so the only uncertainty is TEA’s own scoring, which the agency does not publish an error for.

Provenance. Read from abor_reference.tea_campus_history and tea_campus_current in ALEX’s warehouse on 2026-09-18, rating year 2025-26. The campus report (TAPR) figures in the second table are for 2024-25 — a different year from the ratings, because that is the latest TEA has published. Distance and drive time are routed by the OSRM public routing service, with a straight-line fallback; where a campus could not be located, both are left blank.

Grades and scores are the Texas Education Agency A–F accountability ratings for 2025-26, read from ALEX’s TEA tables. The domain grades are TEA’s three: Student Achievement, School Progress and Closing the Gaps. “Statewide” is the share of Texas campuses at the same level scoring at or above this one in 2025-26. Trend is the change in overall score from 2024-25 to 2025-26; the district trend compares the district’s current and prior ratings. District rank orders the district’s rated campuses at the same level by score, with their average. The campus table is TEA’s Texas Academic Performance Report (TAPR) for 2024-25, the latest year loaded, all students: enrollment, the share economically disadvantaged, students per teacher and average teacher experience, the share of STAAR tests at Meets Grade Level or above across all subjects, attendance and chronic absenteeism, and for high schools the four-year longitudinal graduation rate and the share of graduates enrolled in Texas higher education. Values TEA masks for small groups are shown as a dash. Distance and drive time are routed from the property to the campus location; where a campus could not be located, both are left blank rather than estimated.

Not Market Averages

Measured on this property, not on a market average

Nothing here is a market average. Every figure is computed from this property’s own three comparable sales — how far the method missed when each was hidden and re-predicted, how much the three disagree with one another, how far they had to be moved to match this house, and how old they are. Open either panel below: the first states each measure in plain terms, the second gives the formal treatment.

Click here for the four firmness measures
Hiding each comparable and re-predicting it from the other two, the method is off by about this much. 7.9% Leave-one-out back-test, mean absolute error. Each sale is held out and re-predicted by the rest, re-running all 56 adjustment factors. On $464,529 that is roughly ±$36,700.
How much the three comparables disagree with each other once adjusted to this house. 19.5% Adjusted spread. Highest adjusted comparable against lowest, as a share of the estimate. A wide spread would mean the middle of it was not a better answer than its edges.
How far the engine had to move these houses to make them comparable to this one. ±40.1% average Mean absolute adjustment. Small adjustments mean the comparables were genuinely alike to begin with. The largest single adjustment in this set is $114,876.
On average, these sales closed this long ago. 147 days Mean comparable age. Older sales are carried to today explicitly by the market-timing factor rather than used unchanged.
Click here for full statistics — the four firmness measures

1. Leave-one-out back-test (mean absolute percentage error). For i = 1..n the full pipeline — comparable adjustment across all 56 factors, then adjustment-weighted reconciliation — is refit on D \ {(xᵢ, yᵢ)} and evaluated at xᵢ. The reported statistic is n⁻¹ Σ |ŷ₋ᵢ − yᵢ| / yᵢ. LOOCV is approximately unbiased for expected out-of-sample risk but has high variance: the n training sets share n−2 observations, so fold estimates are strongly positively correlated and the apparent precision is optimistic (Hastie, Tibshirani & Friedman, ESL 2nd ed., §7.10). At n = 3 this is the binding caveat and the reason the figure is presented as descriptive of this comparable set rather than as a population claim.

2. Adjusted spread — a dispersion measure, not an error measure. (max vᵢ − min vᵢ) / V̂ over the adjusted comparable values. It quantifies agreement among estimates, which is not the same quantity as distance from the truth and can move in the opposite direction. Comparables adjusted by one rate table share method and data; their errors are positively correlated, so a tight spread partly reflects common specification rather than convergence on the right answer. Formally, for estimates with pairwise error correlation ρ, the variance of their mean is (σ²/n)[1 + (n−1)ρ] — under strong positive ρ the effective sample size is far below n, and observed spread understates true uncertainty. Low spread with a shared bias is the classic route to a confident wrong answer, which is why this page reports spread and held-out error separately and never substitutes one for the other.

3. Mean absolute adjustment — a proxy for accumulated model error. n⁻¹ Σ gᵢ/Pᵢ, gross adjustment over net sale price. Every adjustment is itself an estimate carrying its own error, so total adjustment magnitude bounds how much of a final figure is model rather than observed transaction. Under the approximation that adjustment errors are independent with variance proportional to magnitude, the error contributed grows with √Σg² — which is the premise the reconciliation weighting rests on, and the reason gross rather than net is the right summary. It is an ordering statistic, not a calibrated error: it ranks comparables by how much work was done to them without asserting how wrong that work was.

4. Mean comparable age — exposure to temporal drift. Mean days between each comparable’s close date and the valuation date. Time adjustment removes the expected market movement, so residual risk is the variance of the index estimate over the carry interval, not the movement itself. Under the Case–Shiller error decomposition that variance grows roughly linearly in the interval, Var ≈ σ²ᵘΔt + 2σ²ᵤ, so age is a proxy for how much index uncertainty each comparable imports. Here the mean is 147 days; the relevant exposure is not the elapsed time but the month-to-month volatility of the index across it, which the market-timing section quantifies directly.

What the four are jointly evidence of. Taken together they describe the provenance of the estimate, not its accuracy: how much of the final figure is observed transaction and how much is model (mean absolute adjustment), whether the independent indications agree (adjusted spread), how much temporal extrapolation was required (comparable age), and how the method performed when it could not see the answer (the back-test). Only the last is an error measure. The first three can all look excellent on a valuation that is wrong.

Why they cannot simply be combined. A reader’s instinct is to average them into a single quality number. They are not independent, and the dependence runs the wrong way. Comparables selected for similarity have small adjustments by construction — the cascade chose them for it — so a low mean adjustment partly reflects the selection rule rather than the market. Tight agreement among comparables adjusted by one rate table reflects the common rate table. Formally, for estimates with pairwise error correlation ρ, the variance of their mean is (σ²/n)[1 + (n−1)ρ]: at ρ near 1 the effective sample size approaches 1 regardless of n, and observed spread understates true uncertainty by an unknown factor. Nothing here estimates ρ — three observations cannot.

The failure mode this section exists to expose. Comparables that are genuinely alike, adjusted by a shared and slightly wrong rate table, will produce a tight spread, small adjustments, and a low back-test error — while every one of them is biased in the same direction. Correlated error is invisible to all four measures simultaneously. That is not a hypothetical on this property: two comparables sit on about double the land, all take a large adjustment from the same land rate, and if that rate is wrong they move together. It is the reason the land section states its rate and its source rather than burying them.

What none of these four is. None is a confidence interval, and none should be read as one. Coverage statements live in Cross Validation, where they are derived from the held-out error distribution rather than from agreement among comparables — agreement among correlated estimates being precisely the thing that can be high while the answer is wrong.

The land

Land Values from County Appraisal District

The comparables sit on 0.51 to 1.02 acres; this property sits on 0.53. Any difference is priced at a rate taken from the target’s own parcel — what Williamson County says its land is worth, divided by how much land the county says it has. Nothing here is estimated or taken from a listing portal.

Baseline= CAD land value ÷ CAD lot size
= $124,254 ÷ 23,086.8 sf  =  $5.3820 / sf
ParcelCAD acresLot sizeCAD land valueCounty $/sf
4319 Miramar Drthe subject, R045978 0.53023,087 sf$124,254$5.38
4301 Verde Vis1.02044,431 sf$192,386$4.33
609 Esparada Dr0.51022,216 sf$120,938$5.44
303 Sequoia Spur1.01043,996 sf$191,196$4.35

The county’s acreage and the MLS lot size agree exactly here — 0.530 acres is 23,086.8 square feet, and the listing says 23,086.8 — so the two independent records of this parcel’s size do not conflict.

Adjustments

ComparableLot sizeDifferenceLand adjustmentShare of net
4301 Verde Vis44,431 sf+21,344 sf−$114,87621.1%
609 Esparada Dr22,216 sf−871 sf+$4,6891.0%
303 Sequoia Spur43,996 sf+20,909 sf−$112,53220.5%

Property tax

The county roll, 2025 to 2026

Williamson CAD values the property at $441,435 for 2026, up 2.9% ($12,359) from $429,076 in 2025. Improvement value moved −$638 and land $12,997. Of the change, 95.3% is land and 4.7% improvement — a land reappraisal. The parcel changed +2.9% against a neighborhood median of +0.9% across 1,048 parcels, a larger change than 77.2% of them. The listing reports annual taxes of $6,675, 1.51% of the county market value.

WCAD20262025Change%
Improvement value$317,181$317,819 −$638 −0.2%
Land value$124,254$111,257 +$12,997 +11.7%
Total market value$441,435$429,076 +$12,359 +2.9%

Against its peers

4319 Miramar Dr (this parcel) +2.9%Neighborhood (1,048) +0.9%ZIP 78628 (24,238) -3.1%Georgetown (62,812) -3.5%Williamson County (286,358) -4.6%

Change in total market value, 2025 to 2026. Bars to the left of the line fell; this parcel is in red.

Peer groupParcelsMedian change
4319 Miramar Drthis parcel, R045978+2.9%
NeighborhoodG150584F - SERENADA (ALL SECTIONS) AND COUNTRY WEST1,048+0.9%
ZIP 7862824,238−3.1%
Georgetown62,812−3.5%
Williamson County286,358−4.6%
Measure, 2026This parcelNeighborhood medianDifference
Land value per acre547 comparable parcels$234,442$190,000+23.4%
Improvement value per square foot1,032 comparable parcels$155.25$163.35−5.0%
Click here for sources and method — the county roll

Values are the Williamson CAD certified market roll for parcel R045978 (SERENADA WEST SEC 4, BLOCK D, LOT 29, ACRES 0.530): land plus improvement equals the total. They are the county’s mass-appraisal figures for taxation, not an opinion of this property’s market value. The roll’s total is not used as a value anywhere in this report. Its land component is used, and only for one purpose: the land section prices lot-size differences at this parcel’s own county land rate, and says so there. Nothing else in the valuation reads from the roll.

“Of the change, X% is land.” Those two shares are of the ABSOLUTE component movements, not of the net change: |Δland| ÷ (|Δland| + |Δimprovement|) and the same for improvement. Here land moved +$12,997 and improvement −$638, which is why the shares can read as 95/5 while the net change is smaller than the land movement alone. No exclusions are applied to the peer groups: splits, new construction and parcels that changed use are all in them, which is part of why a county median can fall while a settled neighbourhood rises.

No interval is attached to any peer median. Each is the median of a complete enumeration of that group’s parcels carrying both rolls — every parcel in the neighbourhood, the ZIP, the city or the county, not a sample of them — so a sampling interval would be answering a question nobody asked. The uncertainty that does exist is in the appraisal itself: the county’s mass-appraisal model, its choice of neighbourhood codes, and any parcel whose use or boundaries changed between rolls. None of that has a published error, and none is invented here.

Peer medians are the median percentage change in total market value between the 2025 and 2026 rolls across every parcel in each group carrying both years. “A larger change than N%” is the share of the neighborhood’s parcels whose change was below this parcel’s (10th–90th percentile −3.2% to +4.8%). Unit values compare land per acre and improvement per square foot with the neighborhood median among parcels carrying those measures. Taxing units: Williamson CAD (CAD); Williamson CO (GWI); Wmsn CO FM/RD (RFM); Georgetown ISD (SGT); Wmsn ESD #8 (F08).

Annual taxes are as reported on the listing; the effective rate divides them by the county market value, before any exemption the owner holds, so a buyer’s bill can differ.

Market timing

How a past sale is carried to today

The factors and adjustments in this section are ALEX’s. Where an agent has changed or zeroed a market-timing line, the adjustment details for that comparable show the agent’s figure and this table still shows ours, so the two can be compared rather than one quietly replacing the other.

A sale that closed months ago is a snapshot in time. To use it as evidence about today, it has to be moved forward by however much the market moved in between. That is what the market timing line on each comparable does.

ComparableClosedIndex thenIndex nowFactorAdjustment
609 Esparada Dr2026-01166.975167.1721.0012+$568
4301 Verde Vis2026-05166.970167.1721.0012+$656
303 Sequoia Spur2026-06171.992167.1720.9720−$15,406

Adjustment = net sale price × (factor − 1). All come from the Williamson County index, built from 19,824 repeat-sale pairs.

What the index numbers mean. They are not dollars and not percentages — they are a scale, and every scale needs a zero. Here the zero is January 2016, set to 100. That month is the base by construction: the regression fixes its value and measures every later month against it. So a reading of 167.172 for September 2026 means Williamson County house prices are 67.2% above where they stood in January 2016.

Click here for full statistics — the index estimator

What no figure in this section carries. No standard error is computed for any monthly index level, and therefore none for the factor applied to a comparable. The repeat-sales regression would support one — the residual variance of the log price differences is estimable — but it is not computed in this pipeline, and an interval that has not been computed is not printed. What the section does state instead is the evidence behind each month: the number of repeat-sale pairs, and the shrinkage weight applied when a month is thin. A month resting on a handful of pairs is pulled toward the metropolitan path for exactly the reason an interval would be wide there.

Specification. Bailey, Muth & Nourse (1963), A Regression Method for Real Estate Price Index Construction, JASA 58(304), 933–942. For a dwelling transacting at periods s and t with s < t, the model is ln(Pᵗ/Pₛ) = βᵗ − βₛ + εₛᵗ. Let X be the N × T design matrix with xᵢₛ = −1, xᵢᵗ = +1 and zero elsewhere. Then β̂ = (X′X)⁻¹X′y with yᵢ = ln(Pᵗ/Pₛ). The base period is normalised to β₀ = 0 and its column deleted; retaining it makes X′X singular, since the indicators sum to zero by construction. The index is Iₜ = 100·exp(β̂ₜ).

Identification. The estimator is consistent for the market factor under the assumption that E[ε | X] = 0 — that idiosyncratic price changes are mean-independent of when a property happened to transact. Dwelling fixed effects difference out exactly because the same unit appears with opposite signs on both sides, which is what removes omitted-variable bias from unobserved quality and is the method’s central advantage over a hedonic index.

Error structure. Case & Shiller (1987, New England Economic Review; 1989, AER 79(1), 125–137) decompose εₛᵗ = (Hᵗ − Hₛ) + (uᵗ − uₛ), where H is a Gaussian random walk in the property’s own value and u is i.i.d. transaction noise. The implied variance Var(ε) = σ²ᵘ(t−s) + 2σ²ᵤ grows linearly in the holding interval, motivating their three-stage weighted least squares: OLS, regress squared residuals on the interval, then GLS with weights 1/√(â + b̂(t−s)). This index does not apply that weighting. It instead excludes holds under 180 days and annualised changes beyond ±60% outright — a trimming rather than a down-weighting strategy, which is more conservative, discards information the WLS would retain, and is stated rather than left implicit.

Shrinkage. County-month estimates are shrunk toward the metropolitan estimate in log space, ln Iᶜₜ = w ln Iᶜₜᶜ + (1−w) ln Iᵐₜ with w = n/(n+k), k = 20. This is the James–Stein / empirical-Bayes posterior mean under a normal-normal hierarchy where k is the ratio of within-month sampling variance to between-county prior variance; here it is fixed rather than estimated. The effect is that thin months borrow strength in proportion to how little evidence they carry of their own — September 2026, with 38 pairs, sits at w = 0.655.

Known limitations. (i) Sample selection. The index is estimated only on dwellings that transacted at least twice, which is not a random sample of the stock; properties that turn over frequently differ systematically, and no Heckman-type correction is applied. (ii) Revision. Estimates for recent periods move as later transactions arrive, so the most recent months are the least stable — a property shared with all repeat-sales indices including S&P CoreLogic Case–Shiller. (iii) Renovation. The dwelling-constant assumption fails where a property was materially improved between sales; the ±60% annualised filter removes the extreme cases and cannot remove the moderate ones. (iv) Aggregation. A single county index assumes a common market factor across submarkets within it.

Empirical note on smoothing. A trailing three-month average of the log index was tested against the raw monthly series by ten-fold cross-validation on 20,631 Williamson County pairs (September 2026), scoring out-of-sample RMSE of predicted log price change. It was worse in 10 of 10 folds (mean difference +0.00171, 95% CI [+0.00139, +0.00202], t = +10.70); a five-month window was worse again (t = +22.77). Monotonicity in the smoothing window is the signature of signal removal rather than noise suppression, consistent with the shrinkage above already performing the variance reduction a second pass would duplicate.

Cross Validation

The error measured on this property

Which figure these measures belong to. Every error, band and interval in this section is computed on ALEX’s own comparable analysis value of $464,529. If an agent has changed an adjustment, the figure at the head of the report moves; these measures do not move with it, because they measure how ALEX’s method performed on this property, not how a revised figure would perform. A changed valuation carries no measured band until the change is tested the same way.

How this was tested, in one sentence: we covered up one of the three sales, worked out what this method would have said that house was worth using only the other two, then uncovered it and measured how close we got — and did that three times, once for each sale.

Three tests is the minimum from which a spread can be quoted. These figures are real and specific to this house, but thin: one unusual comparable moves them materially. Treat them as the shape of the uncertainty, not a precise measurement of it.

Confidence interval 0.0990 the 68% band Two times in three, the true figure should land within ±9.90% — that is ±$46,001, or $418,528 to $510,531. The wider 95% band is 0.1452: ±$67,467, or $397,062 to $531,997.
Typical spread — FSD 12.24% ±$56,863 on this value How much estimates like this one scatter around the truth. It is the same measure Zillow (5%) and HouseCanary (3%) publish for their own models on this property, computed on very different samples.
RMSLE 0.1040 lower is better A single overall error score that treats being 10% too high and 10% too low as equally wrong, and stops one big miss dominating. It has no meaning on its own — it is for comparing one method against another.
Median error 6.82% mean 7.89%, worst 15.38% The middle of the misses — the straightest answer to “how far off is this likely to be”.
4301 Verde Vis +1.5%609 Esparada Dr -15.4%303 Sequoia Spur +6.8%

What happened when each sale was hidden and predicted from the others: how far off that prediction was, as a share of what the sale actually netted. Past 10% is red.

Held outActual netPredicted from the other twoError
4301 Verde Vis$544,200$552,211+1.47%
609 Esparada Dr$482,500$408,294−15.38%
303 Sequoia Spur$549,700$587,201+6.82%
Click here for full statistics — the two ways this report converts a net price to a contract price

Why there are two, and where each is used. This valuation is a net sale price. Two places in the report need a contract price instead, and they do not use the same conversion, so both are stated here rather than left to be discovered.

One: the neighbourhood ratio, used for comparison. ρ̂ = median(net / close) over closed sales in Serenada West (the subject's subdivision) in the last 24 months, n = 16, with a minimum of 15 sales before the level is used at all. Here ρ̂ = 0.9940, an implied seller contribution of 0.60% of the contract price. It grosses our net figure up for the Benchmarks comparison and sets the “how far above the evidence the asking price sits” percentage. It is a median of ratios, unweighted, with no interval reported: at n = 16 a median carries real sampling error, and none is claimed for it.

Two: the pool share, used for the offer ladder. ŝ = Σwᵢ(Cᵢ+Rᵢ)/Pᵢ ÷ Σwᵢ, a recency-weighted MEAN of the actual contribution share over the negotiation pool, with its own drawer in the recommendation section. The two differ in estimator (median of a ratio against a weighted mean of a share), in pool (a subdivision over 24 months against an MLS-area price band over 36), and in what they measure (net divided by close, which includes anything that separated the two, against closing costs and repairs specifically).

What that means for a reader. The two conversions do not have to agree, and on this property they do not: they differ by -0.42 percentage points of the contract price. Neither is used to alter the valuation itself, which is and remains a net sale price. Where a contract price appears, the conversion behind it is named in that section.

Click here for full statistics — how these comparables were selected

Why this matters for everything above. The held-out error, the bands and the confidence score all rest on the comparables actually chosen. If the selection rule were unstated, a reader could not judge whether the held-out comparables and the subject are alike enough for the error measured on them to say anything about the subject. The rule is therefore stated in full.

The cascade. Candidate sales are tested against an ordered list of 453 criteria sets, tightest first. T0 requires the same street and the same subdivision, within 3 years of age, 15% of living area, matching story count, pool and new-construction status, within 180 days, 15% of price per square foot, and a sale-to-list ratio between 0.93 and 1.03. Each later tier relaxes something — first the street, then the subdivision, then the school district, then out to a distance limit, with the last tier at five miles. The engine stops at the first tier that yields at least six qualifying sales, then scores those and keeps the best three. This report stopped at the tier named in the metastrip at the head of the page, and that position is also what feeds the first factor of the evidence-quality score.

Scoring inside the tier. Qualifying sales are ranked by a weighted score: recency 0.25, living-area proximity 0.20, year-built proximity 0.15, price-per-square-foot proximity 0.125, beds 0.05, baths 0.05, lot 0.05, same street 0.05, view proximity 0.03, same story and same pool 0.015 each, walkability 0.015, less a penalty of 0.05 for days on market far from the tier’s median and 0.03 for a stale original list price. These weights are a stated convention, not fitted coefficients, and they decide only the ORDER of sales that have already passed the tier’s hard criteria.

What is removed before any of that. Sales that are not arm’s length (REO, short sale, auction, HUD, corporate-owned, probate, and overbid sales whose remarks name any of those), sales in fair or poor condition, sales in a FEMA flood zone, and the subject’s own earlier sale — that last only when at least three other comparables remain, since below that the set is too thin to refuse the help.

Where the agent stands in this. The comparables on this page were finalised by the agent, so the cascade proposed and a person decided; the tier shown describes where the proposal came from, not a claim that the final three are the cascade’s own.

The limit this creates, stated plainly. Selection is not random. The three sales are the closest matches the rule could find, which is what makes them useful evidence and also what makes the leave-one-out error an optimistic estimate of the error on an arbitrary house: a held-out comparable is being predicted by two sales chosen to be like it. The population figures in this section, measured on a blind holdout of sales the model never saw, are the pessimistic bound on the same question. Both are printed for that reason.

Click here for full statistics — the evidence-quality score

What it scores. 60 of 100 scores the EVIDENCE behind this valuation, not the property and not the probability that the figure is right. It is an index built from four factors, each worth up to 25 points with a floor of 12, because even a weak signal carries some information:

Factor 1, how tight the comparable criteria had to be. The engine walks a cascade of criteria from tightest to loosest and stops at the first that yields a usable set. The score is that stop’s position in the cascade, linearly: the tightest earns 25, the loosest 12. This report stopped at the tier named in the metastrip at the top of the page.

Factor 2, the back-test error. One point deducted per percentage point of leave-one-out error, flooring at 12 once the error reaches 13%. The error itself is in the table above, so this factor is a restatement of measured performance, not a second opinion about it.

Factor 3, how recent the comparables are. One point deducted per 15 days of average age across the set, flooring at 12 once the average passes 195 days.

Factor 4, how much adjustment was needed. One point deducted per percentage point of average absolute adjustment as a share of each comparable’s net price, flooring at 12 at 13%. A set that needed little adjustment scores higher than one rebuilt line by line.

When a factor cannot be measured. It is omitted rather than filled in, and the remaining factors are averaged and scaled to the same 100-point frame; the valuation records which factors were measured. A score built from two real factors is reported as such rather than padded to four with invented inputs.

The tier bonus and the cap. A set drawn from the top of the cascade earns up to 15 additional points, tapering to zero across the top 3% of tiers; the total is then capped at 100. Labels: high at 70 or above, medium at 50 or above, low below 50 — this property scores 60, medium.

What it is not. The deductions above are conventions chosen so the score moves in the right direction and can be read off the inputs; they are not fitted coefficients, the score has no sampling distribution, and no interval is attached to it. It is not a probability and must not be read as one. The only uncertainty this report claims is the held-out error and the bands built from it, immediately above.

Click here for full statistics — accuracy standards and cross-validation methodology

How the population figures were produced. They come from the training run’s own artifact, not from this page: a single temporal split, train on everything closing on or before 2026-03-08 and test on the 6 months after it, so no sale in the test set was seen in any form during fitting. 19,486 test sales out of the full table. Leases are removed, and the metrics are reported segmented by property type and price band as well as overall, because a single blended figure hides exactly the segments that matter. No interval is attached to these population figures in the artifact, and none is invented here.

What they do and do not describe. They describe the ALEX statistical model across a metro-wide holdout. They are not the accuracy of the comparable-sales figure on this page, which is measured by leaving each comparable out, immediately above; the two are different estimators on different samples and are reported separately for that reason.

Part one — what a valuation should be measured against

This section is reserved. The standard is being supplied separately and will be placed here in full. It is deliberately empty rather than filled with a plausible-looking threshold: a benchmark that this report happened to pass would be indistinguishable, to a reader, from one it had actually been measured against. Until the paper is here, no accuracy standard is being claimed or implied anywhere in this report.

What is known today, pending that standard. Three sets of numbers exist and none of them is a standard — they are measurements, and they are not measuring the same thing:

1. This property. Median error 6.82%, FSD 12.24%, RMSLE 0.1040, 68% interval ±9.90%, from three held-out comparables. Specific to this house, and thin: three observations describe this comparable set accurately and generalise weakly.

2. This model, across its blind holdout. MdAPE 7.77%, MAPE 11.22%, RMSLE 0.1605 over 19,486 sales the model never saw during training. This is the right basis for any general claim about the method, and it is deliberately not what the tiles above report — those answer “how accurate is this valuation”, which is a different question.

3. What the vendors publish for themselves. HouseCanary: national MdAPE 2.8% over 1,994,203 transactions internally, 2.9% on a blind third-party test. Zillow: median error 1.79% on-market and 7.20% off-market. Two cautions before either is treated as a bar to clear. Both are national aggregates dominated by dense, homogeneous tract housing where an automated model does best, and neither is conditioned on the kind of property this is. And the on-market figure is not independent of the asking price — both vendors consume list price, which is the whole point made in the Benchmarks section.

What a standard would need to specify to be usable here, and what the forthcoming paper will presumably settle: which statistic is the criterion (MdAPE, FSD, coverage, or a hit rate); whether it is a point threshold or a distribution; whether it applies per property or per portfolio; how it is conditioned on price band, property type and evidence depth; and whether an interval must be calibrated — that is, whether a stated 68% interval must actually contain the outcome 68% of the time, which is a materially stricter requirement than a median error threshold and the one most AVM disclosures avoid.

Where this report already exceeds common practice, whatever the standard turns out to be: the comparables are named, every adjustment is itemised with its reason, the estimator and its weights are stated, and the held-out test results are printed so every summary statistic can be recomputed by hand. A standard can be applied to a number; it can only be verified against a number whose derivation is visible.

Part two — how this valuation was measured

Why hold anything out at all. Scoring a method on the same sales it was given measures how well it can interpolate data it has already seen, which is not the quantity anyone cares about. The quantity of interest is expected loss on an observation drawn from the same population and not used in fitting: Err = E(X,Y)[L(Y, f̂(X))]. Resubstitution error is downward-biased for it by construction. The land difference is where this bites on this property. Two of the three comparables sit on about double its land, so each time one was held out the method had to price that extra land with little recent evidence of its own to price it from.

Design: leave-one-out cross-validation. For i = 1,…,n the entire pipeline — all 56 adjustment factors, the repeat-sales time carry, and the inverse-gross-adjustment reconciliation — is refit on D−i = D \ {(xᵢ,yᵢ)} and evaluated at xᵢ, giving CV = n⁻¹ Σᵢ L(yᵢ, f̂−i(xᵢ)). Critically, the held-out comparable is excluded from the reconciliation as well as from the adjustment grid; leaving it in the weighting while removing it from the grid would leak the answer through the weights.

Properties, and the one that governs here. LOOCV is approximately unbiased for Err — it trains on n−1 of n points, so the pessimistic bias of k-fold at small k is largely absent. It pays for this in variance: the n training sets are nearly identical, sharing n−2 observations, so the fold-level estimates are strongly positively correlated and Var(CV) does not contract at the 1/n rate an independent average would. At n = 3 this is the binding constraint on everything below (Hastie, Tibshirani & Friedman, ESL 2nd ed., §7.10–7.12; Efron & Tibshirani, An Introduction to the Bootstrap, 1993, ch. 17).

Forecast standard deviation. FSD = sd{ln(ŷᵢ/yᵢ)}, computed on the held-out pairs. The log ratio is scale-invariant, so the statistic is comparable across price points and across vendors without normalisation. Under an approximately log-normal error, exp(±FSD) bounds a one-standard-deviation multiplicative interval; the log-normality is an approximation and is not tested at n = 3. This is the AVM industry's disclosure convention, which is what makes the comparison against the vendors’ published figures like-for-like in definition — though not in sample, since theirs are national aggregates over millions of transactions and this is three comparables for one house.

Interval construction, and why not a normal-theory interval. The 68% and 95% figures are empirical quantiles of the absolute percentage error distribution {|ŷᵢ−yᵢ|/yᵢ}, not μ̂ ± zσ̂. A normal-theory interval would import a distributional assumption the sample cannot support and would be symmetric in levels, which a multiplicative error process is not. The trade-off is explicit: a quantile at the 95th percentile of a handful of observations is an extrapolation of the fitted distribution, not an observed exceedance rate, and is labelled as such wherever it is quoted. Coverage validity requires exchangeability between the held-out comparables and the subject — the assumption underpinning conformal prediction (Vovk, Gammerman & Shafer, Algorithmic Learning in a Random World, 2005) — which close, same-era selection supports without guaranteeing. It is worth noting that conformal width calibration was measured for this pipeline and rejected: a six-month-forward holdout violates exchangeability outright, and coverage did not improve.

RMSLE. RMSLE = [n⁻¹ Σ (ln(1+ŷᵢ) − ln(1+yᵢ))²]1/2. Squaring in log space penalises proportional rather than absolute error and bounds the leverage of a single large residual relative to RMSE. It is asymmetric in levels: under-prediction is penalised more heavily than over-prediction of the same absolute size, since |ln(1−δ)| > ln(1+δ). The 1+ offset is vestigial at these magnitudes. RMSLE is a relative statistic with no absolute interpretation — it exists to rank methods, and quoting it alone says nothing.

MdAPE and MAPE. MAPE = n⁻¹ Σ |ŷᵢ−yᵢ|/yᵢ; MdAPE is the median of the same summands. MAPE is undefined at y = 0 and structurally asymmetric: unbounded above for over-prediction, bounded by 1 for under-prediction, so it systematically favours methods that under-forecast. The median is reported alongside because it has a 50% breakdown point against the mean's 1/n. Here 7.89% against 6.82% indicates no single dominating residual.

Hit rates. PPEk = n⁻¹ Σ 1{|ŷᵢ−yᵢ|/yᵢ ≤ k} at k = 0.10 and 0.20 — the empirical CDF of absolute percentage error at two points. On this property: 2 of 3 within 10%, 3 of 3 within 20%. Reported as counts, not percentages: the estimator takes only 4 values at n = 3, and a Wilson score interval on 2/3 spans roughly [0.21, 0.94].

What is not claimed, stated plainly. No standard error is attached to any statistic here. The sampling distribution of a standard deviation at this sample size is severely right-skewed, and a nonparametric bootstrap resamples the same few points — it would produce an interval, and that interval would be an artefact of the resampling rather than evidence about the population. No hypothesis is tested and no p-value is reported, because there is no null here worth rejecting at this sample size. These figures are offered as descriptive statistics of this comparable set, which is what they are and the most that three observations can support. The corresponding population-level figures — MdAPE 7.77%, RMSLE 0.1605 over a blind holdout of 19,486 sales — exist and are the right basis for any claim about the method in general; they are not what this page reports, because the numbers asked for are the ones specific to this property.

What others say

Benchmarks

ALEX — this report, shown for comparison $467,352 Our net $464,529 grossed back up to a contract price using the ratio Serenada West sales actually run at, so it can be set beside the outside estimates below. Three Serenada West sales, each adjusted to this property across 56 factors and reconciled by adjustment-weighted mean; supported range $411,170–$501,746.

ALEX, grossed to a contract price $467,352HouseCanary $552,812Zillow $526,800WCAD 2026 roll $441,435

Each estimate as a point, its published range as the line through it. Ours is red.

The outside estimates below were not used to produce our number. Each “vs ALEX” percentage compares that vendor with $467,352 — our net figure grossed to a contract price by the neighbourhood ratio described in the drawer below — not with the figure at the head of this report, because the vendors publish contract-price estimates and a net price is not the same quantity. They are here so you can see what else is being said about this house. Each is a contract price, and each is compared with the ALEX figure above.

Outside sourceValueRangevs ALEXWhat it is
HouseCanaryconfidence 97%, HIGH $552,812$534,631–$570,993 +18.3% A national automated valuation, quoting its own forecast standard deviation of 3%. Retrieved with this valuation run.
Zillow Zestimate $526,800$500,460–$553,140 +12.7% Zillow rates its own confidence here “Good”, forecast standard deviation 5%. It cannot see the interior, the condition, or the subdivision boundary that governs this comparable set.
Williamson CADparcel R045978 $441,435 −5.5% 2026 appraised value: land $124,254 plus improvement $317,181; prior year $429,076. A tax assessment, not an opinion of market value — produced en masse for taxation, lagging by design, and capped by statute.
Cotalityformerly CoreLogic pending Reserved — contract renewal pending. No valuation is retrieved until the agreement is executed. This is a commercial position, not a gap in this property’s data and not a failure of the model. Nothing is estimated in its place, and this row will carry a real figure when access resumes.

There is a floor, around ±5%. No model can reliably predict a single sale much closer, because sales themselves disagree — the same house listed twice in one month will not fetch the same price. Different buyers, different negotiation, different day. That scatter is how houses are sold, not a flaw in any model, and more data will not remove it. Treat any accuracy claim much tighter than ±5% on one property with suspicion.

What we will and will not claim about that floor. The ±5% figure above is a judgement about transaction noise, not a published result we can cite, and this report does not rest on it: nothing here is measured against 5% or against any industry target. What is measurable is on this page — the held-out error on this property and the population figures in Cross Validation — and that is the only standard this valuation asks to be judged against.

The vendor models are proprietary. Zillow, HouseCanary and Cotality — the company that acquired CoreLogic — publish neither their methods, their parameters nor their data, and their accuracy figures are self-scored on samples they choose. Nobody outside those companies can reproduce a number, test an interval, or prove any of it wrong — which is the minimum before a figure counts as a defensible statistic. Good commercial products; not evidence in the sense this report uses the word.

Now the thing worth noticing. HouseCanary lands 2.4% above the $540,000 asking price, Zillow lands 2.4% below the $540,000 asking price. This report is 13.5% below it. The reason is simple: they use the asking price; we do not. HouseCanary’s own brief lists MLS “listed prices and contract prices” as inputs and says a new listing triggers an immediate re-valuation. Zillow publishes two accuracy figures for the same reason: 1.79% median error on-market against 7.20% off.

Set that 1.79% against the floor. It is about three times tighter than transaction noise alone permits — the signature of a model that already knows the asking price, not one that has out-predicted the market. So when an automated estimate hugs an asking price, much of what you are seeing is that asking price reflected back. To HouseCanary, Zillow and Cotality, a list price is real information about what a seller and their listing agent want. It is not independent evidence of value, which is what this report was asked for. Ours uses only what three neighbouring sales actually sold for.

This does not make the vendor numbers useless. We use them as benchmarks. We do not include them in our valuation model.

Click here for full statistics — what each vendor publishes, and what it does not

Why an undisclosed model is not a defensible statistic. A figure is defensible when someone else can reproduce it: the estimator is specified, the data are identified, and the validation protocol is stated. Both vendors publish outputs and accuracy summaries; neither publishes the estimator, the trained parameters, or the data. Their accuracy figures are therefore self-reported and self-scored on samples they select. This is normal commercial practice and not a criticism of their competence — it is a statement about what an outside reader can check, which is: the claim, but not the calculation.

Zillow, in its own words. The Zestimate is described as combining “public records, MLS data and user-submitted home details into Zillow’s proprietary home valuation model,” offered as “a transparent, free starting point,” and explicitly “not an appraisal and can’t be used in place of an appraisal.” Published accuracy: nationwide median error 1.79% on-market and 7.20% off-market — meaning half of on-market Zestimates fall within 1.79% of the eventual sale price. Zillow states directly that on-market estimates are more accurate “because more up-to-date information is available, including listing details and recent market activity.” No specification, feature list, or training data is published.1

HouseCanary, from its published technical brief. Considerably more is disclosed. The algorithm runs in three stages: (1) query and clean data; (2) build localised price indices; (3) train machine-learning models on time-adjusted historical prices. Models are fitted at census-tract level, borrowing from neighbouring tracts where a tract is too thin to model alone, with neighbours used for training but excluded from accuracy testing. Two index models are built — median price and median price per square foot — at census-block level, and all historical sales and list prices are carried to current values through them. The response variable is each price expressed as a percent deviation from its block’s current median for that property type; several ML models per tract are then fitted to explain that deviation and combined into one estimate. Inputs include 3,100+ county assessors, 2,700+ county recorders over 20 years, MLS characteristics, listed prices and contract prices, mortgage balances and distress measures. In non-disclosure states such as Texas, MLS contract prices substitute for recorded sale prices where an arm’s-length sale can be jointly verified — directly relevant here.2

HouseCanary’s FSD, and an important caveat they state themselves. FSD is trained on the census-tract empirical error distribution and depends both on that spread and on how much the component models disagree for the individual property. The confidence score is simply 1 − FSD, which is why the 97% confidence reported for this property corresponds to an FSD of 0.03. The interval is constructed for approximately 68% coverage: P(1 ± FSD). And in their own words: “We assume symmetry but not normality... using 2*FSD to estimate a new interval is not guaranteed to yield an approximate 95% coverage probability.” Anyone doubling their FSD to get a 95% band is doing something the vendor explicitly warns against.

What HouseCanary validates, and how. Monthly internal testing on a rolling six-month window, plus quarterly blind third-party testing. As of July 2019: national MdAPE 2.8% over 1,994,203 transactions internally, and 2.9% on the third-party blind sample for Q1 2019. They report hit rate, MdAPE, median signed error (as a bias check), within-5/10/20%, and the share of sales falling inside their own 68% interval — a coverage check, which is the right diagnostic and more than most vendors publish. Results are available by state and MSA on request, which is the boundary of the disclosure: the validation design is public, the validation data are not.

The comparison this page makes, and its limit. FSD is a shared convention, so this report’s 12.24% and theirs are the same quantity — the standard deviation of log(estimate ÷ outcome). But they are computed on different samples by different methods: theirs across millions of national transactions, this one across three comparables for this house. A national median error says what happens to a typical home; it does not say what happens to this particular property and its own comparables. Reading a vendor figure as a promise about this house would be a category error, and reading ours as vendor-grade inaccuracy would be the same error in reverse.

What this report exposes by contrast. The estimator is stated, the comparables are named with MLS numbers, every adjustment is itemised with its reason, the index method is a published 1963 regression with its exclusions listed, and the held-out test results are printed so every summary statistic can be recomputed by hand. Whether the answer is right is a separate question; whether it can be checked is not.

1 zillow.com/z/zestimate — quotations from Zillow’s published “What is a Zestimate?” page. · 2 HouseCanary Valuation Technical Brief, © 2019 HouseCanary Inc., read in full. Figures are as of that publication and may have moved since. · 3 cotality.com (formerly CoreLogic) — no valuation estimate was obtained; the contract is pending renewal.

Editing copy — sheet and report. Click any wording and type over it. Figures, charts and controls are locked. Nothing here saves by itself — press Save when you are done.