Stability of cross-validation and minmax-optimal number of folds

نویسندگان

  • Ning Xu
  • Jian Hong
  • Timothy C. G. Fisher
چکیده

In this paper, we analyze the properties of cross-validation from the perspective of the stability, that is, the difference between the training error and the error of the selected model applied to any other finite sample. In both the i.i.d. and non-i.i.d. cases, we derive the upper bounds of the one-round and average test error, referred to as the one-round/convoluted Rademacher-bounds, to quantify the stability of model evaluation for cross-validation. We show that the convoluted Rademacher-bounds quantify the stability of the out-of-sample performance of the model in terms of its training error. We also show that the stability of cross-validation is closely related to the sample sizes of the training and test sets, the ‘heaviness’ in the tails of the loss distribution, and the Rademacher complexity of the model class. Using the convoluted Rademacher-bounds, we also define the minmax-optimal number of folds, at which the performance of the selected model on new-coming samples is most stable, for cross-validation. The minmax-optimal number of folds also reveals that, given sample size, stability maximization (or upper bound minimization) may help to quantify optimality in hyper-parameter tuning or other learning tasks with large variation.

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Near-Optimal Bounds for Cross-Validation via Loss Stability

Multi-fold cross-validation is an established practice to estimate the error rate of a learning algorithm. Quantifying the variance reduction gains due to cross-validation has been challenging due to the inherent correlations introduced by the folds. In this work we introduce a new and weak measure called loss stability and relate the cross-validation performance to this measure; we also establ...

متن کامل

Cross-Validation and Mean-Square Stability

k-fold cross validation is a popular practical method to get a good estimate of the error rate of a learning algorithm. Here, the set of examples is first partitioned into k equal-sized folds. Each fold acts as a test set for evaluating the hypothesis learned on the other k − 1 folds. The average error across the k hypotheses is used as an estimate of the error rate. Although widely used, espec...

متن کامل

An a Priori Exponential Tail Bound for k-Folds Cross-Validation

We consider a priori generalization bounds developed in terms of cross-validation estimates and the stability of learners. In particular, we first derive an exponential Efron-Stein type tail inequality for the concentration of a general function of n independent random variables. Next, under some reasonable notion of stability, we use this exponential tail bound to analyze the concentration of ...

متن کامل

Determining optimal value of the shape parameter $c$ in RBF for unequal distances topographical points by Cross-Validation algorithm

Several radial basis function based methods contain a free shape parameter which has  a crucial role in the accuracy of the methods. Performance evaluation of this parameter in different  functions with various data has always been a topic of study. In the present paper, we consider studying the methods which determine an optimal value for the shape parameter in interpolations of radial basis  ...

متن کامل

A Study of Cross - Validation and Bootstrapfor Accuracy Estimation and Model

We review accuracy estimation methods and compare the two most common methods: cross-validation and bootstrap. Recent experimental results on artiicial data and theoretical results in restricted settings have shown that for selecting a good classiier from a set of classi-ers (model selection), tenfold cross-validation may be better than the more expensive leave-one-out cross-validation. We repo...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

عنوان ژورنال:
  • CoRR

دوره abs/1705.07349  شماره 

صفحات  -

تاریخ انتشار 2017