Showing posts with label model. Show all posts
Showing posts with label model. Show all posts

Friday, March 9, 2012

How linear regression choose his regressor ?

I would like to understand the algorithm that the linear regression method uses to choose the regressors in the model from a list of possible regressors.

I think that it is different from the common methods used in statistics like stepwise, forward or backward.

Laura Lerner

When you build the model using the wizard, all continuous input columns are flagged as regressors.

You can check this by looking at the Modeling Flags property of each input column in the designer.

(edit) Jamie points out to me that the modeling flags only indicate potential regressors. The actual regressors are determined by the internals of the algorithm. We have no public documentation available on that process, but we can aim to provide more detail in future docs or whitepapers.

|||

I would like to know how the algorithm works, whether there are parameters I can change to influence the results.

Laura Lerner

|||

Laura Lerner wrote:

I would like to know how the algorithm works, whether there are parameters I can change to influence the results.

Laura Lerner

If you wish to make a specific input column a regressor for linear regression, you can do that by changing the Modeling Flag property of that column to Regressor.

I am wondering if there is a specific reason that you would want to influence the choice of regressors rather than explicitly setting them. That would be good to know, to inform our future plans.

Thanks

|||

For my purposes I would like to see the p values associated with the regressors so that I can decide which variables to include or not based on them, myself. Furthermore, in order not to run into a problem of multicollinearity it would be useful to see a correlation matrix as part of regression analysis and computed Variance Inflation Factors would definitely help as well. If we are providinga wish listSmile, also at least a graph of samples of residuals would be great to have.

Best Regards,

|||

Thanks! I am quite happy to take wishes. Whether or when I can grant them is a different matter, but the more information we have on what people want - and, very importantly, why - the better.

I have had a number of requests for computed Variable Inflation Factors.

Feel free to contact me offline donald_dot_farmer_at_microsoft_dot_com if you want to discuss potential features and business cases in more detail.

Thanks

|||You can use the FORCE_REGRESSOR parameter to guarantee that the algorithm will use a particular regressor, regardless of the algorithm used to determine applicability. The target of FORCE_REGRESSOR does have to be marked as a regressor though.|||

I am building several automatic linear regressions using code and not the interface.

I give to the model hundreds variables it can choose and the result is a linear regression with very few variables.

I would like first to understand how it chooses the variables and whether there are parameters that I can use to influence the number of variables the linear regression will choose.

I used for many year SAS and there I know that I can use parameters like the significance level, number of variables to include, etc...

Laura Lerner

How is the 'Score' value derived in the Lift chart/Mining legend for Data Mining Models?

Hi,

I have just run a simple data set through a model to predict a simple true or false value (i.e. binary output)

The Lift Chart/Mining Legend in Analysis Services shows three results – Score, Population Correct (%), and Predict Probability (%)

Population Correct I beleive is the percentage of predictions it got right out of the total number of predictions it tried to make. Is this correct?

However, I can’t work out how the other two are derived in particular the 'SCORE'. To give a live example the scores were as follows:

Model Score Pop Correct Pred Probability

Decision Trees 0.83 76.59% 54.28%

Neural Network 0.75 67.63% 50.05%

Ideal Model 100.00%

Can anyone help with this and give a detailed explanation?

Many thanks,

S Rajput

Hi

The Predict Probability is the probability of the most popular prediction state for the model.

The Score is the log scatter score for a scatter plot of (x=Actual Value,y=Predicted Value) where each point has an associated probability. There are a lot of examples on the mathematical series for this calculation which you can look up.

Hope this helps.

Shuvro

|||

So does a larger score mean a better fit for the model? How is a scatter score calculated?

|||

Here're some more details about the scatter score calculation:

This score is the (geometric) mean score of all the points constituting the scatter plot.

Here is how it works.

0) Each point on a scatter plot corresponds to a test case and has the form (a,b(M)), where a is the actual attribute value for the case and b(M) is the value predicted using the model M;

1) First define the score(a,b(M)) for *one* individual scatter point; to do this compare the (a,b(M)) to the best prediction we can do without using any model; that prediction is of course marginalMean; so

score(a,b(M)) = likelihood ( b(M), given a, given M) / likelihood( marginalMean, given a)

2) Then average out all the scores across the entire scatter plot;

Technically is was simpler to average them out as

score = ( Product[ score(a,b) | forall (a,b) in the scatter plot])^(1/n), where n is the number of points in the scatter plot.

So that’s how it is currently done.

Details:

The statistical meaning of the individual point score is the predictive lift for the model M measured at that point. It is essentially the same notion as the Predict Likelihood fraction for the case contributed by the continuous attribute for which the scatter plot has been built.

The specific formula for this is score(a,b) = pdf( N(a, predictStdev), b) / pdf( N(a, marginalStdev), marginalMean)

where N(mu, sigma) is the normal distribution with mean of mu and Stdev sigma and

pdf(N(mu, sigma), x) is its probability density function described by exp((x-mu)^2/(2*sigma))/sqrt(2Pi)/sigma.

Wednesday, March 7, 2012

How is the 'Score' value derived in the Lift chart/Mining legend for Data Mining Model

Hi,

I have just run a simple data set through a model to predict a simple true or false value (i.e. binary output)

The Lift Chart/Mining Legend in Analysis Services shows three results – Score, Population Correct (%), and Predict Probability (%)

Population Correct I beleive is the percentage of predictions it got right out of the total number of predictions it tried to make. Is this correct?

However, I can’t work out how the other two are derived in particular the 'SCORE'. To give a live example the scores were as follows:

Model Score Pop Correct Pred Probability

Decision Trees 0.83 76.59% 54.28%

Neural Network 0.75 67.63% 50.05%

Ideal Model 100.00%

Can anyone help with this and give a detailed explanation?

Many thanks,

S Rajput

Hi

The Predict Probability is the probability of the most popular prediction state for the model.

The Score is the log scatter score for a scatter plot of (x=Actual Value,y=Predicted Value) where each point has an associated probability. There are a lot of examples on the mathematical series for this calculation which you can look up.

Hope this helps.

Shuvro

|||

So does a larger score mean a better fit for the model? How is a scatter score calculated?

|||

Here're some more details about the scatter score calculation:

This score is the (geometric) mean score of all the points constituting the scatter plot.

Here is how it works.

0) Each point on a scatter plot corresponds to a test case and has the form (a,b(M)), where a is the actual attribute value for the case and b(M) is the value predicted using the model M;

1) First define the score(a,b(M)) for *one* individual scatter point; to do this compare the (a,b(M)) to the best prediction we can do without using any model; that prediction is of course marginalMean; so

score(a,b(M)) = likelihood ( b(M), given a, given M) / likelihood( marginalMean, given a)

2) Then average out all the scores across the entire scatter plot;

Technically is was simpler to average them out as

score = ( Product[ score(a,b) | forall (a,b) in the scatter plot])^(1/n), where n is the number of points in the scatter plot.

So that’s how it is currently done.

Details:

The statistical meaning of the individual point score is the predictive lift for the model M measured at that point. It is essentially the same notion as the Predict Likelihood fraction for the case contributed by the continuous attribute for which the scatter plot has been built.

The specific formula for this is score(a,b) = pdf( N(a, predictStdev), b) / pdf( N(a, marginalStdev), marginalMean)

where N(mu, sigma) is the normal distribution with mean of mu and Stdev sigma and

pdf(N(mu, sigma), x) is its probability density function described by exp((x-mu)^2/(2*sigma))/sqrt(2Pi)/sigma.