SPB Git

spb/wp5_uqo Public

UQO Working Paper No. 5 — Airbnb, residential rents and housing market pressure.

TeX 53.4% Python 46.5%
12.5 KB · 91 lines latex
Raw Blame History
1% Author: Simon-Pierre Boucher — contact@spboucher.ai2% =============================================================================3% 04_methodology.tex4% =============================================================================5\section{Methodology}\label{sec:methodology}67This section presents the econometric models employed to quantify the association between Airbnb activity and residential rents, discusses identification, and describes the machine-learning robustness framework.89\subsection{Model 1: Hedonic Rent Model with Airbnb Exposure}\label{sec:model1}1011Our baseline specification is a hedonic rent equation \citep{rosen1974hedonic, malpezzi2003hedonic} in which the log of monthly rent is regressed on a measure of nearby Airbnb activity and a vector of dwelling-level controls:1213\begin{equation}\label{eq:hedonic_rent}14    \ln(\text{rent}_i) = \alpha + \beta \cdot \texttt{airbnb\_count}_{i,r} + \mathbf{X}_i' \boldsymbol{\gamma} + \sum_{c} \delta_c \cdot \mathbf{1}[\text{city}_i = c] + \varepsilon_i,15\end{equation}1617\noindent where $\text{rent}_i$ is the monthly rent of listing $i$; $\texttt{airbnb\_count}_{i,r}$ is the number of Airbnb listings within buffer radius $r$ of listing $i$ (our preferred specification uses $r = 500$\,m); $\mathbf{X}_i$ is a vector of dwelling characteristics including the number of bedrooms, number of bathrooms, building type (Apartment, House, Row/Townhouse), and interior size (where available); $\delta_c$ denotes city fixed effects that absorb all time-invariant, city-level unobservables (including average neighbourhood quality, local labour-market conditions, and municipal regulations); and $\varepsilon_i$ is a mean-zero error term.1819The coefficient of interest, $\beta$, has a semi-elasticity interpretation: it measures the approximate percentage change in monthly rent associated with one additional Airbnb listing within the buffer, conditional on observed dwelling characteristics and city.  A positive and statistically significant $\hat{\beta}$ is consistent with the hypothesis that Airbnb activity is associated with higher residential rents, though it does not, by itself, establish causality.  The reliance on spatial fixed effects and multiple functional forms of the exposure variable follows the specification guidance of \citet{kuminoff2010which}, whose simulation evidence identifies these as the most effective safeguards against omitted-variable bias in cross-sectional hedonic models.2021We estimate Equation~\eqref{eq:hedonic_rent} by ordinary least squares (OLS) with heteroskedasticity-robust standard errors (HC1).2223\subsection{Model 2: Hedonic Airbnb Pricing Model}\label{sec:model2}2425To examine the pricing determinants of Airbnb listings and the relationship between local rents and short-term rental prices, we estimate a hedonic model of Airbnb nightly prices:2627\begin{equation}\label{eq:hedonic_airbnb}28    \ln(\text{price}_j) = \alpha' + \theta \cdot \overline{\text{rent}}_c + \mathbf{Z}_j' \boldsymbol{\lambda} + \sum_{c} \delta'_c \cdot \mathbf{1}[\text{city}_j = c] + u_j,29\end{equation}3031\noindent where $\text{price}_j$ is the nightly price of Airbnb listing $j$; $\overline{\text{rent}}_c$ is the mean log monthly rent in city $c$ (obtained from the rental dataset); $\mathbf{Z}_j$ includes listing characteristics---the number of bedrooms and bathrooms, guest capacity, amenities count, superhost status, and the entire-home indicator; and $u_j$ is an error term.3233The coefficient $\theta$ captures the association between the local residential rent level and Airbnb pricing.  A positive $\hat{\theta}$ would suggest that Airbnb hosts in higher-rent cities charge higher nightly prices, consistent with the opportunity-cost channel: when the forgone rental income from listing a unit on Airbnb is higher, hosts set higher nightly prices to compensate.  Note that when city fixed effects are included, $\overline{\text{rent}}_c$ is collinear with the city dummies; in such specifications, $\theta$ is identified from the city-level variation absorbed by the fixed effects, and we report results both with and without city fixed effects.3435\subsection{Model 3: City-Level Interaction Model}\label{sec:model3}3637To characterise the bidirectional relationship between Airbnb activity and rents at the city level, we estimate a forward and a reverse aggregate regression:3839\begin{align}40    \overline{\ln(\text{rent})}_c &= a_1 + b_1 \cdot \text{airbnb\_count}_c + \mathbf{W}_c' \boldsymbol{\phi}_1 + e_{1c}, \label{eq:city_rent} \\[6pt]41    \text{airbnb\_count}_c &= a_2 + b_2 \cdot \overline{\ln(\text{rent})}_c + \mathbf{W}_c' \boldsymbol{\phi}_2 + e_{2c}, \label{eq:city_airbnb}42\end{align}4344\noindent where overlines denote city-level means, $\text{airbnb\_count}_c$ is the total number of Airbnb listings in city $c$, and $\mathbf{W}_c$ contains city-level mean bedrooms and bathrooms.  The coefficients $b_1$ and $b_2$ describe the cross-city co-movement of rents and Airbnb activity.  Because many of the 153 cities in the aggregated sample contribute few underlying listings, these regressions should be interpreted as descriptive rather than inferential.4546\subsection{Model 4: Spatial Analysis}\label{sec:model4}4748To account for spatial dependence in rents, we estimate a spatial autoregressive (SAR) model and a spatial error model (SEM) alongside the OLS baseline.  The SAR model augments the hedonic equation with a spatial lag of the dependent variable:4950\begin{equation}\label{eq:spatial_lag}51    \ln(\text{rent}_i) = \alpha'' + \rho \cdot \sum_{k} w_{ik} \ln(\text{rent}_k) + \beta' \cdot \texttt{airbnb\_count}_{i,r} + \mathbf{X}_i' \boldsymbol{\gamma}' + \sum_{c} \delta''_c \cdot \mathbf{1}[\text{city}_i = c] + \eta_i,52\end{equation}5354\noindent where $\mathbf{W} = [w_{ik}]$ is a row-standardised $k$-nearest-neighbour spatial weights matrix with $k = 5$, so that $\sum_k w_{ik} \ln(\text{rent}_k)$ is the mean log rent of the five nearest rental listings.  The spatial autoregressive parameter $\rho$ captures the degree to which rents co-move within a spatial neighbourhood, after controlling for observed characteristics.  The SEM instead places the spatial process in the disturbance, $\eta_i = \lambda \sum_k w_{ik} \eta_k + \nu_i$, capturing spatially correlated unobservables.  Both models are estimated by the feasible generalised-moments procedures of \citet{kelejian1998generalized} and \citet{kelejian1999generalized} (\texttt{GM\_Lag} and \texttt{GM\_Error}), which avoid the simultaneity bias that OLS estimation of Equation~\eqref{eq:spatial_lag} would entail; see \citet{anselin1988spatial} and \citet{lesage2009introduction} for the underlying theory.5556This specification serves two purposes.  First, the spatial models absorb variation from spatially correlated unobservables (e.g., neighbourhood amenities that affect both rents and Airbnb desirability).  Second, comparing $\hat{\beta}'$ from Equation~\eqref{eq:spatial_lag} with $\hat{\beta}$ from Equation~\eqref{eq:hedonic_rent} provides a diagnostic for the sensitivity of the Airbnb coefficient to spatial confounders.5758In addition to the spatial lag model, we examine the robustness of the Airbnb exposure measure by varying the buffer radius $r$ across 250\,m, 500\,m, 1\,km, and 2\,km.  This exercise traces out the spatial decay of the association: if the Airbnb--rent relationship is genuinely local, we expect $\hat{\beta}$ to be largest at small radii and to attenuate as the buffer expands.5960\subsection{Model 5: Quantile Regression}\label{sec:model5}6162To investigate heterogeneity in the Airbnb--rent association across the conditional rent distribution, we estimate quantile regressions \citep{koenker1978regression, koenker2001quantile}:6364\begin{equation}\label{eq:quantile}65    Q_{\tau}\!\left[\ln(\text{rent}_i) \mid \mathbf{X}_i, \texttt{airbnb\_count}_{i,r}\right] = \alpha_\tau + \beta_\tau \cdot \texttt{airbnb\_count}_{i,r} + \mathbf{X}_i' \boldsymbol{\gamma}_\tau + \sum_{c} \delta_{c,\tau} \cdot \mathbf{1}[\text{city}_i = c],66\end{equation}6768\noindent for quantile indices $\tau \in \{0.10, 0.25, 0.50, 0.75, 0.90\}$.  City fixed effects are restricted to the five largest cities (with the remainder grouped) to keep the quantile estimation well conditioned.  The quantile-specific coefficient $\beta_\tau$ measures the marginal association between Airbnb exposure and the $\tau$th quantile of log rent.  If the effect of Airbnb is larger at the upper tail of the distribution, this may indicate that short-term rental activity disproportionately affects higher-quality or higher-rent segments of the market.  Standard errors are the asymptotic kernel-based estimates of \citet{koenker1978regression} as implemented in \texttt{statsmodels}.6970\subsection{Model 6: Machine-Learning Robustness}\label{sec:model6}7172We complement the parametric analysis with four machine-learning methods: the LASSO \citep{tibshirani1996regression}, the elastic net \citep{zou2005regularization}, random forests \citep{breiman2001random}, and gradient boosting \citep{friedman2001greedy}, alongside an OLS benchmark; all are implemented in scikit-learn \citep{pedregosa2011scikit}.  These models are trained to predict $\ln(\text{rent}_i)$ from the Airbnb exposure measures at the 500\,m radius (count, density, mean price, entire-home share), dwelling characteristics (bedrooms, bathrooms), and geographic coordinates, and are evaluated on a held-out test set (20\% of the sample) using root mean squared error (RMSE), mean absolute error (MAE), and $R^2$.7374The machine-learning models serve three purposes.  First, they provide a benchmark for the predictive accuracy of the hedonic model: if the OLS model achieves comparable $R^2$ to the flexible ML models, this suggests that the linear specification is not severely misspecified.  Second, LASSO and elastic net coefficients reveal which variables are selected as predictors, providing a data-driven assessment of the importance of Airbnb exposure relative to dwelling characteristics.  Third, random forest and gradient boosting models yield variable-importance measures (SHAP values) that quantify the contribution of each feature to predictive accuracy, without imposing functional-form assumptions.7576\subsection{Identification and Endogeneity}\label{sec:identification}7778We are explicit about the identification challenges inherent in our cross-sectional design.  The primary concern is that Airbnb listing density is endogenous to local rent levels and neighbourhood desirability.  Three channels of endogeneity are relevant:7980\begin{enumerate}[label=(\roman*)]81    \item \textbf{Reverse causality:} High rents may attract Airbnb hosts (because the opportunity cost of leaving a unit vacant is high, and short-term rental income can offset high carrying costs), creating a positive bias in the estimated $\hat{\beta}$.82    \item \textbf{Omitted variables:} Neighbourhood amenities---such as proximity to cultural attractions, restaurants, and transit---that are unobserved in our data may simultaneously attract Airbnb guests and drive up residential rents, again biasing $\hat{\beta}$ upward.83    \item \textbf{Spatial sorting:} Airbnb hosts may select locations based on unobserved characteristics (e.g., architectural charm, noise tolerance norms) that also affect residential rents.84\end{enumerate}8586Our city fixed effects control for all time-invariant, city-level confounders, including differences in local housing-market tightness, tourism infrastructure, and regulatory regimes.  Within-city variation in Airbnb density is our identifying variation, and this variation is plausibly driven by both locational amenities (tourist attractiveness) and the supply of suitable housing units---neither of which is fully exogenous.  We therefore interpret our estimates as conditional correlations rather than causal effects.  The inclusion of spatial lags, robustness checks across buffer radii, and machine-learning diagnostics provides indirect evidence on the plausibility and stability of our estimates, but a definitive causal statement would require either panel data with suitable fixed effects or a credible instrument for Airbnb supply---neither of which is available in our setting.8788\subsection{Standard Errors}8990Throughout the analysis, we report heteroskedasticity-robust standard errors (HC1) as our baseline; Section~\ref{sec:robustness} re-computes the preferred specification with standard errors clustered at the city level.  For the quantile regressions, we report the asymptotic standard errors produced by the kernel-based estimator in \texttt{statsmodels}.  For the machine-learning models, we report held-out test-set performance metrics, with regularisation parameters selected by 5-fold cross-validation.91