Team Ai
Datasetpublic

cive202/Lifestyle_Preferences_Dataset

🏠 Lifestyle Preferences Dataset A tabular dataset for roommate compatibility, preference modeling, and recommendation systems cive202/Lifestyle_Preferences_Dataset Β· Hugging Face Hub Overview The Lifestyle Preferences Dataset is a tabular dataset containing lifestyle and preference information for machine learning, data analysis, and recommendation-system applications. It is well suited to studying relationships between individual lifestyle preferences… See the full description on the dataset page: https://huggingface.co/datasets/cive202/Lifestyle_Preferences_Dataset.

sourceHugging Faceupdated 2mo agoView on Hugging Face
1likes21downloads
Dataset Card

<div align="center">

🏠 Lifestyle Preferences Dataset

A tabular dataset for roommate compatibility, preference modeling, and recommendation systems

cive202/Lifestyle_Preferences_Dataset Β· Hugging Face Hub

![License: MIT](https://opensource.org/licenses/MIT) Task Size Format

</div>


Table of Contents


Overview

The Lifestyle Preferences Dataset is a tabular dataset containing lifestyle and preference information for machine learning, data analysis, and recommendation-system applications.

It is well suited to studying relationships between individual lifestyle preferences and compatibility-related outcomes β€” particularly in roommate matching, preference modeling, personalization, and recommendation systems.

πŸ’‘ Example research questions - Which lifestyle preferences are most strongly associated with compatibility? - Can users be grouped by similar lifestyle patterns? - Can a model predict compatibility between two individuals? - Which preferences matter most when recommending potential roommates? - How can preference-based matching systems be designed?

Files

FileDescriptionSize
roommate_dataset.csvPrimary dataset~1.79 MB
README.mdDataset documentation and metadataβ€”
Note: The repository documentation does not yet specify the complete column schema. Inspect the CSV header and dtypes before modeling β€” see Loading the Dataset below.

Intended Use

  • β€”Roommate compatibility analysis
  • β€”Lifestyle preference modeling
  • β€”Recommendation systems
  • β€”User similarity and clustering
  • β€”Exploratory data analysis
  • β€”Tabular machine-learning experiments
  • β€”Preference-based matching systems
  • β€”Educational projects involving real-world-style tabular data

Loading the Dataset

Option 1 β€” pandas

python
import pandas as pd

df = pd.read_csv("roommate_dataset.csv")

print(df.head())
print(df.shape)
print(df.columns.tolist())
print(df.info())

Option 2 β€” πŸ€— Datasets library

python
from datasets import load_dataset

dataset = load_dataset(
    "csv",
    data_files="https://huggingface.co/datasets/cive202/Lifestyle_Preferences_Dataset/resolve/main/roommate_dataset.csv"
)

print(dataset)

Recommended Preprocessing

Before using the dataset for machine learning, consider the following checklist:

  • β€”[ ] Check for missing values
  • β€”[ ] Check for duplicate records
  • β€”[ ] Inspect categorical variables and their possible values
  • β€”[ ] Encode categorical variables where required
  • β€”[ ] Scale numerical features where appropriate
  • β€”[ ] Check class balance if using a target variable
  • β€”[ ] Separate training and evaluation data before fitting models
  • β€”[ ] Screen for potentially sensitive or personally identifying information

Evaluation

<table> <tr><th>Task Type</th><th>Suggested Metrics</th></tr> <tr><td><b>Classification / Compatibility</b></td><td>Accuracy Β· Precision Β· Recall Β· F1 Β· ROC-AUC Β· Confusion Matrix</td></tr> <tr><td><b>Regression</b></td><td>MAE Β· MSE Β· RMSE Β· RΒ²</td></tr> <tr><td><b>Recommendation / Matching</b></td><td>Precision@K Β· Recall@K Β· NDCG</td></tr> </table>

The appropriate methodology depends on the specific target variable and task definition.


Data Splits

No official train/validation/test split is currently documented for this dataset.

If you create your own split:

  • β€”Avoid leakage between related individuals or records
  • β€”Document your random seed and splitting strategy for reproducibility

Data Collection

The repository does not currently provide detailed documentation about the original data-collection process, sampling methodology, population, or annotation procedure.

⚠️ Do not assume the dataset is representative of any particular geographic, demographic, or real-world population without additional information from the dataset creator.

Privacy & Sensitive Information

Lifestyle and roommate-preference data can potentially reflect individuals' habits, routines, or personal preferences.

  • β€”Review the actual dataset contents before deploying models built on this data
  • β€”Avoid attempting to identify individuals represented in the dataset
  • β€”Handle any personally identifying or sensitive attributes in accordance with applicable privacy and data-protection requirements

Bias & Limitations

  • β€”Unknown sampling methodology
  • β€”Potential sampling or demographic bias
  • β€”Limited documentation of data provenance
  • β€”Unknown representativeness of the underlying population
  • β€”Possible subjective interpretation of lifestyle preferences
  • β€”Possible mismatch between collected/synthetic preferences and real-world roommate behavior

Models trained on this dataset should not be automatically assumed to generalize to broader populations.


License

Released under the MIT License, as specified in the Hugging Face repository metadata.

Verify separately that the underlying data itself is appropriate for your intended use, and that any third-party data is used consistently with its original licensing and terms.


Contact

Maintained by cive202 on the Hugging Face Hub.

For questions, corrections, or additional documentation, use the Community tab on the dataset repository or reach out via the maintainer's Hugging Face profile.