Why it matters
Part Two makes the case for a rental property data strategy, including the compelling reasons why it’s a good idea to do something about bad data rather than living with it.
If you stood in a room of real estate professionals and declared that rental property data is a mess, it’s doubtful anyone would challenge you. Rental property data problems are the industry’s open secret. What is less obvious is how exactly we got here. That is the question you have to answer before any of it can be fixed, and it is where this three-part series on the messy reality of rental property data begins.
It’s no secret that the industry struggles with data issues. Plenty of large industries struggle with technology adoption, but real estate seems even more susceptible to the challenge than most. When you look for the reasons, they don’t start at the property level. They start at the top, at a very high level.
One reason is that this is a fragmented market, historically held by a relatively small number of individuals who had no compelling incentive to push enterprise-level systems. If you own a handful of assets and know them personally, a spreadsheet and a good memory feel sufficient. Multiply that decision across decades of ownership and the industry never develops a shared expectation of what a system should produce.
On top of that sit many systemic and structural reasons engrained in the industry that have perpetuated the problems with technology use and adoption. These factors have led to a genuinely dismal state for the data those technology systems produce. Rental property data problems start there, not in any one office. The systems themselves are not the villain. What they were asked to do, and what nobody asked them to do, is.
Dig into the details behind why so many rental property data problems exist and the same six causes come up again and again. They are cumulative. Each one on its own is manageable; together they make a portfolio’s data unusable without heavy manual repair.
Properties generate an enormous amount of data, but most of it is generated and stored in various systems, and those systems don’t always talk to each other. Even where APIs exist, the communication back and forth between systems generally has no centralized place to be safely stored. The result is volume without accumulation: a lot of data is produced and very little of it lands anywhere it can be used.
Data was originally entered by humans, in various formats. That inconsistency, combined with the absence of a standard data format, is what prevents machines from reading and processing the data easily. Every field entered by hand without a controlled format is a field that will need to be interpreted later by someone who wasn’t there.
The structure of the data doesn’t focus on residents. It’s about the units, which reflects the previously transactional nature of rental housing. That framing shows up everywhere downstream. It’s easy to ask a system what a unit did last year and much harder to ask what a resident experienced, because nothing was designed around the second question.
Changes in staff or management companies mean that over time, different people are entering the same information into different systems. That produces inconsistencies between reports from various sources that rely on those same data sets. It is exactly why ownership groups so often receive numbers that don’t reconcile against each other.
Property sales happen. Even if a single management company held one standard format across every property it managed, that consistency would likely break as soon as a new ownership group or management company came in. The historical record of an asset is therefore stitched together from however many conventions have governed it.
The rental real estate market doesn’t have a common data structure. Without an industrywide commitment to adopt one data standard, companies find it challenging to scale their data collection and standardization at all. So most ownership groups and property management firms build their own systems. Consequently, they may be collecting information on the same properties in entirely different formats.
The costs of this are easy to underestimate because they rarely appear as a line item. They appear as time, doubt, and decisions made a little later than they should have been.
Each of these rental property data problems is a tax on the same activity: using what you already own to make a decision. The data exists. It is the format, not the volume, that is failing.
The honest answer is that nobody chose this. A fragmented ownership base with no incentive to standardize, decades of manual entry, a unit-centric data structure inherited from a transactional business, and constant churn in both staff and ownership together produced an industry where the same property can be described accurately in several mutually incompatible ways.
That also tells you what a fix for rental property data problems has to look like. It cannot depend on everyone agreeing to use the same system, because the market structure that created the problem is the same market structure that makes universal agreement unlikely. What can work is a layer that accepts data in whatever format each system produces it and converts it into one consistent structure. That is the job of multifamily ETL and data standardization. The mechanics of how raw property exports become comparable records are covered in our walkthrough of the standardization process.
What it does require is deciding, once, what each field means, and then enforcing that decision on everything that arrives. In practice that means a written definition for each field, one accepted format for each value, and a structure that can describe the resident as well as the unit. Those three decisions address the six causes directly. Decentralized data gets one destination. Human entry gets a controlled format. The transactional structure gets a resident dimension. Staff, management, and ownership changes stop resetting the record, because the definition outlives the people. The causes listed above don’t disappear. They stop mattering downstream.
Now you know how we got to the land of bad data. The next two parts take on the questions that follow from it.
Part Two makes the case for a rental property data strategy, including the compelling reasons why it’s a good idea to do something about bad data rather than living with it.
Part Three covers how to actually fix rental property data problems, and why standardization is the practical path rather than industrywide system consolidation.
Rental property data problems begin with market structure. Real estate is a fragmented market historically held by individuals who had no compelling incentive to push enterprise-level systems, and a range of systemic and structural factors engrained in the industry have perpetuated poor technology adoption. Those factors produced a dismal state for the data that technology systems generate: an enormous amount of it, spread across systems that don’t reliably talk to each other.
Data was originally entered by humans in various formats, with no standard format to enforce consistency. Changes in staff and in management companies mean different people enter the same information into different systems over time, and property sales bring new ownership groups and managers with their own conventions. Reports drawn from those different sources then disagree with each other.
The rental real estate market never made an industrywide commitment to adopt one data standard, which makes it challenging for any single company to scale its data collection and standardization. In the absence of a shared structure, most ownership groups and property management firms developed their own systems. They may therefore be collecting information on the same properties in completely different formats.
The causes of messy rental property data are structural, which is exactly why a standardization layer that accepts every system’s format and outputs one consistent structure is the practical way through.
See plans and pricing