What We Learned in Decision 2012

Now the election in US is over. What differs the most from the last presidential election in 2008 is the impacts of new technologies such as blogs and social media. Interestingly, Nate Silver made a surprisingly accurate prediction of the election results in his famous FiveThirtyEight blog. His near-perfect forecast solidifies the relevance and significance of big data solutions. In my opinion, 4 key factors are the critical enablers to unlock the value of big data: Modeling, Algorithm, Statistics, and Semantics (MASS).
  1. Modeling: First and foremost, a good model must be established to represent, capture, and ingest the large amount of data in a structured, unstructured, or semi-structured format. The nature of the data elements is a largely deciding factor for an appropriate data store.
  2. Algorithm: A sophisticated algorithm has to be developed to process the data in an optimized way. An easy-to-use coding method is needed to balance the local processing and global computation in a distributed fashion. For example, historical data can be tapped for generating valuable recommendations based on a user profile by means of the click-through rate and interest match metrics.
  3. Statistics: Statistical data analysis is becoming increasingly important, and open source packages like R make data mining more transparent. Growing commercial supports for R from the major vendors fuel the adoption, integration, and distribution of R.
  4. Semantics: Context-awareness is a must. Simple analysis is no longer sufficient for today's business. Complex analytics requires advanced techniques such as patterns and probabilistic reasoning. Vagueness is inevitable and got be dealt with properly to extract insights from massive data in a fuzzy way.
