Q9Data Preparation for Machine Learning (ML)
HOTSPOT - An ML engineer is working on an ML model to predict the prices of similarly sized homes. The model will base predictions on several features The ML engineer will use the following feature engineering techniques to estimate the prices of the homes: • Feature splitting • Logarithmic transformation • One-hot encoding • Standardized distribution Select the correct feature engineering techniques for the following list of features. Each feature engineering technique should be selected one time or not at all (Select three.)
← → navigate · a answer
Discussion · 7
7
The building size is a standardized distribution
7
Size of building (Square feet or Square Meters) = Logarithmic transformation
Explanation: Building size is a numerical feature that often has a skewed distribution and can have a non-linear relationship with price. Logarithmic transformation works well because:
It helps normalize skewed distributions
It can help linearize the relationship between size and price
It's especially useful for features that follow exponential or multiplicative patterns
Real estate data often shows log-normal distributions
6
city = one-hot encoding; type_year = feature splitting; size of the building = standardized distribution
5
city = one-hot encoding; type_year = feature splitting; size of the building = standardized distribution
4
City (Name) = One-hot encoding Explanation: City names are categorical variables that don't have any numerical relationship with each other. One-hot encoding is the best option for this kind of data because:
It creates binary columns for each unique city
It avoids adding an artificial ordering between cities
It lets the model treat each city as an independent feature
4
Since the question highlights "similarly sized homes", that means those numbers for build size won't show a skewed distribution. So, Size of the building should be standard distribution.
3
Type_year (type of home and year it was built) = Feature splitting Explanation: This feature includes two different pieces of information (type and year) combined in one column. Feature splitting is the right choice because:
It separates the compound feature into its parts
The type can then be one-hot encoded
The year can be handled as a numerical feature
This split lets the model learn from each component separately