Predict House Prices
Problem statement
You are given preprocessed train_data, validation_data, and test_data. Use the labeled training and validation rows to predict a price for every test row.
For this deterministic practice version:
- Do not use
house_idorpriceas model features. - Combine the training and validation rows. Standardize each feature with that combined labeled mean and population standard deviation; ignore features whose labeled standard deviation is
0. - For each test row, select the three nearest labeled rows by squared Euclidean distance, or all labeled rows when fewer than three exist. Break equal-distance ties by labeled input order.
- Transform neighbor prices with
log(1 + price), average those transformed values, apply the inverse transform, and round the predicted price to six decimal places.
Return exactly one column named price, preserving the original test_data row order.
Table schema
Pandas
Use the same input data with any supported language. Open the Schema tab in the editor to see the generated SQL setup or Pandas DataFrames.
train_data
The labeled training split.
| Column | Type | Nullable | Description |
|---|---|---|---|
| house_idPK | Integer | No | — |
| bedrooms | Decimal | No | — |
| bathrooms | Decimal | No | — |
| sqft_living | Decimal | No | — |
| sqft_lot | Decimal | No | — |
| floors | Decimal | No | — |
| waterfront | Integer | No | — |
| age | Decimal | No | — |
| num_views | Decimal | No | — |
| condition | Integer | No | — |
| price | Decimal | Yes | — |
validation_data
The labeled validation split.
| Column | Type | Nullable | Description |
|---|---|---|---|
| house_idPK | Integer | No | — |
| bedrooms | Decimal | No | — |
| bathrooms | Decimal | No | — |
| sqft_living | Decimal | No | — |
| sqft_lot | Decimal | No | — |
| floors | Decimal | No | — |
| waterfront | Integer | No | — |
| age | Decimal | No | — |
| num_views | Decimal | No | — |
| condition | Integer | No | — |
| price | Decimal | Yes | — |
test_data
The unlabeled test split to predict.
| Column | Type | Nullable | Description |
|---|---|---|---|
| house_idPK | Integer | No | — |
| bedrooms | Decimal | No | — |
| bathrooms | Decimal | No | — |
| sqft_living | Decimal | No | — |
| sqft_lot | Decimal | No | — |
| floors | Decimal | No | — |
| waterfront | Integer | No | — |
| age | Decimal | No | — |
| num_views | Decimal | No | — |
| condition | Integer | No | — |
| price | Decimal | Yes | — |
Expected result
Your query or function must return these columns.
| Column | Type | Nullable | Description |
|---|---|---|---|
| price | Decimal | No | — |
Row order: must match exactly. Numeric tolerance: 0.000001.
Constraints
- Every input table is non-empty.
- Training and validation prices are non-null and nonnegative.
- Test prices are null and must not be used.
- Every feature value is finite.
- All three tables have the same feature columns.
- The output has exactly one row per test row.
Source note: Original Capital One data-science assessment task (carousel slide 3).