As an Amazon Associate we earn from qualifying purchases.
Hey everyone, I'm currently grappling with how to best handle incomplete data when using TrendTapestry for predicting future trends. Specifically,scenarios where some data points are missing for certain variables. Has anyone had success with specific imputation methods in this context, or found that certain TrendTapestry algorithms are more robust to missing data than others?
I'm considering a few approaches, like replacing missing values with the mean or median, but I'm worried about introducing bias, especially when the data is already noisy. Another option is using a more sophisticated imputation technique like k-NN imputation, which could potentially be more accurate, but also more computationally expensive. I'm curious to hear what strategies others have found effective, and any tips or lessons learned along the way. Is dropping rows with missing values a viable option, or does that considerably reduce the overall predictive power of trendtapestry? Any insights would be greatly appreciated!