Twitter Geolocation and Regional Classification via Sparse Coding
نویسندگان
چکیده
We present a data-driven approach for Twitter geolocation and regional classification. Our method is based on sparse coding and dictionary learning, an unsupervised method popular in computer vision and pattern recognition. Through a series of optimization steps that integrate information from both feature and raw spaces, and enhancements such as PCA whitening, feature augmentation, and voting-based grid selection, we lower geolocation errors and improve classification accuracy from previously known results on the GEOTEXT dataset.
منابع مشابه
Kernel Density Estimation for Text-Based Geolocation
Text-based geolocation classifiers often operate with a grid-based view of the world. Predicting document location of origin based on text content on a geodesic grid is computationally attractive since many standard methods for supervised document classification carry over unchanged to geolocation in the form of predicting a most probable grid cell for a document. However, the grid-based approa...
متن کاملImage Classification via Sparse Representation and Subspace Alignment
Image representation is a crucial problem in image processing where there exist many low-level representations of image, i.e., SIFT, HOG and so on. But there is a missing link across low-level and high-level semantic representations. In fact, traditional machine learning approaches, e.g., non-negative matrix factorization, sparse representation and principle component analysis are employed to d...
متن کاملFace Recognition using an Affine Sparse Coding approach
Sparse coding is an unsupervised method which learns a set of over-complete bases to represent data such as image and video. Sparse coding has increasing attraction for image classification applications in recent years. But in the cases where we have some similar images from different classes, such as face recognition applications, different images may be classified into the same class, and hen...
متن کاملMicroblog Geolocation using Language Variation Deep Learning
1 David Zucker 2 Department of Computer Science 3 Stanford University 4 Stanford, CA 94305 5 [email protected] 6 7 8 Introduction 9 This experiment investigates the feasibility of geographically locating Twitter users based solely 10 on tweet content through the identification of geographic regional language and dialect patterns. 11 Currently, fewer than 3% of current tweets are configured t...
متن کاملRice Classification and Quality Detection Based on Sparse Coding Technique
Classification of various rice types and determination of its quality is a major issue in the scientific and commercial fields associated with modern agriculture. In recent years, various image processing techniques are used to identify different types of agricultural products. There are also various color and texture-based features in order to achieve the desired results in this area. In this ...
متن کامل