For the tax column, I’d separate an unusually high ongoing burden from any property-specific issue mentioned in the marketing information. Buyers may react differently, and combining them could make the apparent pattern stronger than it is.
Would price per square metre help identify the ambitious listings? It would not replace condition or exact location, but it could show whether the long-lived properties were simply positioned above nearby alternatives in the same broad segment.
The R$4,682,000–R$7,022,000 band is wide enough for buyer expectations to shift within it. I’d split it into smaller bands before concluding there is one 78-day market, while keeping the full-range result visible for comparison.
Agreed, although smaller bands can become noisy if the sample is limited. I’d prioritise completed versus active status and neighbourhood first, then subdivide by price only where there are enough comparable properties to make the result meaningful.
Neighbourhood labels themselves can be misleading near boundaries. If two listings are physically close but assigned different area names, strict labels may separate genuine comparables. A map-based grouping would be useful, without pretending every nearby property is equivalent.
For ambiguous removals, the safest practical step is to ask whoever supplied the status whether it means sold, withdrawn or merely no longer advertised there. If that cannot be established, label it unknown. Forcing a sale outcome would contaminate the comparison.
There is also a correlation problem with flood-risk flags. If flagged properties cluster geographically, the measured delay could reflect location preference, building type or condition. Compare within a narrow area before attributing the difference to flood concern alone.
Seller behaviour may be visible in the sequence: launch price, time to first reduction, size of later reductions, then removal. A seller who makes several changes is sending a different signal from one who holds the same price for months.
I wouldn’t filter out financed buyers just to get a cleaner completion timeline. That would answer a narrower question. Better to preserve them and mark, where known, whether financing may have extended the period after agreement.
So the practical order seems to be: clean reposts, define first-observed dates, split neighbourhoods, classify status, then examine price cuts, condition, tax and flood risk. Only after that would I compare this month with an earlier cohort.
One caution on that order: condition and duplicate identification may need to happen together. A renovated and newly photographed property could be a relisting, but it may also be materially different from its previous version. Link the records without assuming they are unchanged.
If completed-sale information is sparse, could accepted-offer or pending status be used as an interim signal? It would be less final, but potentially closer to the buyer’s decision date than the eventual completed transaction.
Yes, as a separate status rather than a completed sale. Pending transactions can still change, and the meaning of status labels varies by provider. They are useful for timing, but should not be merged into the final-sale group.
I’d now describe the initial result narrowly: active properties in this price range have an observed age around 78 days. That wording avoids claiming they all need 78 days to sell while the completed and withdrawn cohorts are still unresolved.
And if the figure is an average, publish the median beside it. A gap between them would immediately reveal whether a few old listings are driving the headline. No extra causal story is needed until the distribution is visible.
For the monthly comparison, keep the method frozen. Changing neighbourhood boundaries, duplicate rules or condition categories midway could look like a market movement. Record any unavoidable change and calculate the earlier month again under the same definitions.
The strongest answer will probably be a range rather than one number: time to offer where known, time to completed sale, and age at withdrawal or last observation. That preserves the uncertainty while showing whether the 78-day active sample is broadly consistent with actual outcomes.