SYS-CON MEDIA Authors: Liz McMillan, Yeshim Deniz, Elizabeth White, William Schmarzo, Dana Gardner

Related Topics: SYS-CON MEDIA, Artificial Intelligence, @CloudExpo, @DXWorldExpo, @ThingsExpo

SYS-CON MEDIA: Blog Post

Demystifying Data Science | @CloudExpo @Schmarzo #BigData #AI #DataScience #ArtificialIntelligence

Data science is about identifying those variables and metrics that might be better predictors of performance

[Opening Scene]: Billy Dean is pacing the office. He’s struggling to keep his delivery trucks at full capacity and on the road. Random breakdowns, unexpected employee absences, and unscheduled truck maintenance are impacting bookings, revenues and ultimately customer satisfaction. He keeps hearing from his business customers how they are leveraging data science to improve their business operations. Billy Dean starts to wonder if data science can help him. As he contemplates what data science can do for him, he slowly drifts off to sleep, and visions of Data Science starts dancing in his head…

[Poof! Suddenly Wizard Wei appears]: Hi, I’m your data science wizard to help alleviate your data science concerns. I don’t understand why folks try to make the data science discussion complicated. Let’s start simple with a simple definition of data science:

Data science is about identifying those variables and metrics that might be better predictors of performance

The key to a successful analytical model is having a robust set of variables against which to test for their predictive capabilities. And the key to having a robust set of variables from which to test is to get the business users engaged early in the process.

[A confused Billy Dean]: Okay, but I’m still confused. I mean, how does this really apply to my business?

[A patient Wizard Wei]: Well, let’s say that you are trying to predict which of your routes are likely to have under-capacity loads so that you can combine loads. In order to identify those variables that might be better predictors of under-capacity routes, you might ask your business users:

What data might you want to have in order to predict under-capacity routes?

The business users are likely to come up with a wide variety of variables, including:

Customer name Ship to location Customer industry
Building permits Customer tenure Change in customer size
Customer stock price Customer D&B rating Types of products hauled
Time of year Seasonality/Holidays Day of week
Traffic Weather Local Events
Distance from distribution center Open headcount on Indeed.com Tenure of logistics manager

The Data Science team will then gather these variables, perform some data transformations and enrichment, and then look for variables and combinations of variables that yield the best predictive results regarding under-capacity routes (see Figure 1).

Figure 1: Data Science Process

Role of Artificial Intelligence
[A less confuse Billy Dean]:
Ah, I think I understand, but what about all this talk about artificial intelligence? From some of these commercials on TV, it appears that robots with artificial intelligence will be ruling the world. Can you say Skynet?

[A still patient Wizard Wei]: Ah, that’s just marketing. Artificial intelligence is just one of many different tools in the predictive analytics kit bag of a data scientist. But artificial intelligence – while embracing some very sophisticated mathematical, data enrichment and computing techniques – is really pretty straightforward. All artificial intelligence is trying to do is to find and quantify relationships between variables buried in large data sets (see Figure 2).

Figure 2: Understanding Artificial Intelligence

[An inquisitive Billy Dean]: Okay, I’m starting to get it, but there seems to be some many
different analytic and predictive algorithms from which to choose. How does the business user know where to start?

[A growing frustrated Wizard Wei]: Ah, that’s the secret to the process. Business users don’t need to know which algorithms to use; they need to be able to identify those variables that might be better predictors of performance. It is up to the data science team to determine which variables are the most appropriate by testing the different algorithms.

Data Mining, Machine Learning and Artificial Intelligence (including areas such as cognitive computing, statistics, neural networks, text analytics, video analytics, etc.) are all members of the broader category of data science tools. Our data scientist team has experts in each of these areas, though no one data scientist is an expert at all of them (in spite of what they tell me). The different data science tools are used in different scenarios for different needs. Think of one of your mechanics. They have a large toolbox full of different tools. They determine what tools to use to fix a truck based upon the problem they are trying to solve. That’s exactly what a data scientist is doing, just with a different toolbox of algorithms.

No single algorithm is best over whole domain; so different algorithms are needed to cover different domains. Often combinations of algorithms are used in order to achieve the best results. To be honest, it’s like a giant jigsaw puzzle with the data science team constantly testing different combinations of metrics, data enrichment and algorithms until they find the combination that yields the best results.

[An enlightened Billy Dean]: I think I’ve finally got it. All of these different algorithms and techniques are just trying to help predict what is likely to happen so that I can make better operational and customer issues. But what’s the realm of what’s possible with data and analytics; I mean, how effective can my organization become at leveraging data and analytics to power my business?

[A proud Wizard Wei]: Great question, and the heart of the big data and data science conversation. Figure 3 shows how you could use these different data science tools to progress up the Big Data Business Model Maturity Index; to transition from running your business on Descriptive analytics that tell you what happened (Monitoring stage) to Predictive analytics that tell you what is likely to happen (Insights stage) to Prescriptive analytics that tell you what they should do (Optimization stage).

Figure 3: Leveraging Artificial Intelligence to drive Business Value

In the end, the data and the analytics are only useful if they help you optimize key operational processes, reduce compliance and security risks, uncover new revenue opportunities and create a more compelling, more prescriptive customer engagement. In the end, data and analytics are all about your business.

[A satisfied Billy Dean]: That’s great Wizard Wei! Thanks for your help!

Now, what can you do about my taxes…

To learn more about “Demystifying Data Science”, come to my Dell EMC World session: “Demystifying Data Science: A Pragmatic Guide To Building Big Data Use Cases” See you there!!

The post Demystifying Data Science appeared first on InFocus Blog | Dell EMC Services.

Read the original blog entry...

More Stories By William Schmarzo

Bill Schmarzo, author of “Big Data: Understanding How Data Powers Big Business” and “Big Data MBA: Driving Business Strategies with Data Science”, is responsible for setting strategy and defining the Big Data service offerings for Hitachi Vantara as CTO, IoT and Analytics.

Previously, as a CTO within Dell EMC’s 2,000+ person consulting organization, he works with organizations to identify where and how to start their big data journeys. He’s written white papers, is an avid blogger and is a frequent speaker on the use of Big Data and data science to power an organization’s key business initiatives. He is a University of San Francisco School of Management (SOM) Executive Fellow where he teaches the “Big Data MBA” course. Bill also just completed a research paper on “Determining The Economic Value of Data”. Onalytica recently ranked Bill as #4 Big Data Influencer worldwide.

Bill has over three decades of experience in data warehousing, BI and analytics. Bill authored the Vision Workshop methodology that links an organization’s strategic business initiatives with their supporting data and analytic requirements. Bill serves on the City of San Jose’s Technology Innovation Board, and on the faculties of The Data Warehouse Institute and Strata.

Previously, Bill was vice president of Analytics at Yahoo where he was responsible for the development of Yahoo’s Advertiser and Website analytics products, including the delivery of “actionable insights” through a holistic user experience. Before that, Bill oversaw the Analytic Applications business unit at Business Objects, including the development, marketing and sales of their industry-defining analytic applications.

Bill holds a Masters Business Administration from University of Iowa and a Bachelor of Science degree in Mathematics, Computer Science and Business Administration from Coe College.

Latest Stories
Eric Taylor, a former hacker, reveals what he's learned about cybersecurity. Taylor's life as a hacker began when he was just 12 years old and playing video games at home. Russian hackers are notorious for their hacking skills, but one American says he hacked a Russian cyber gang at just 15 years old. The government eventually caught up with Taylor and he pleaded guilty to posting the personal information on the internet, among other charges. Eric Taylor, who went by the nickname Cosmo...
Most modern computer languages embed a lot of metadata in their application. We show how this goldmine of data from a runtime environment like production or staging can be used to increase profits. Adi conceptualized the Crosscode platform after spending over 25 years working for large enterprise companies like HP, Cisco, IBM, UHG and personally experiencing the challenges that prevent companies from quickly making changes to their technology, due to the complexity of their enterprise. An accomp...
DevOpsSUMMIT at CloudEXPO, to be held June 25-26, 2019 at the Santa Clara Convention Center in Santa Clara, CA – announces that its Call for Papers is open. Born out of proven success in agile development, cloud computing, and process automation, DevOps is a macro trend you cannot afford to miss. From showcase success stories from early adopters and web-scale businesses, DevOps is expanding to organizations of all sizes, including the world's largest enterprises – and delivering real results. Am...
The benefits of automated cloud deployments for speed, reliability and security are undeniable. The cornerstone of this approach, immutable deployment, promotes the idea of continuously rolling safe, stable images instead of trying to keep up with managing a fixed pool of virtual or physical machines. In this talk, we'll explore the immutable infrastructure pattern and how to use continuous deployment and continuous integration (CI/CD) process to build and manage server images for any platfo...
Nicolas Fierro is CEO of MIMIR Blockchain Solutions. He is a programmer, technologist, and operations dev who has worked with Ethereum and blockchain since 2014. His knowledge in blockchain dates to when he performed dev ops services to the Ethereum Foundation as one the privileged few developers to work with the original core team in Switzerland.
It cannot be overseen or regulated by any one administrator, like a government or bank. Currently, there is no government regulation on them which also means there is no government safeguards over them. Although many are looking at Bitcoin to put money into, it would be wise to proceed with caution. Regular central banks are watching it and deciding whether or not to make them illegal (Criminalize them) and therefore make them worthless and eliminate them as competition. ICOs (Initial Coin Offer...
Business professionals no longer wonder if they'll migrate to the cloud; it's now a matter of when. The cloud environment has proved to be a major force in transitioning to an agile business model that enables quick decisions and fast implementation that solidify customer relationships. And when the cloud is combined with the power of cognitive computing, it drives innovation and transformation that achieves astounding competitive advantage.
René Bostic is the Technical VP of the IBM Cloud Unit in North America. Enjoying her career with IBM during the modern millennial technological era, she is an expert in cloud computing, DevOps and emerging cloud technologies such as Blockchain. Her strengths and core competencies include a proven record of accomplishments in consensus building at all levels to assess, plan, and implement enterprise and cloud computing solutions. René is a member of the Society of Women Engineers (SWE) and a m...
The current environment of Continuous Disruption requires companies to transform how they work and how they engineer their products. Transformations are notoriously hard to execute, yet many companies have succeeded. What can we learn from them? Can we produce a blueprint for a transformation? This presentation will cover several distinct approaches that companies take to achieve transformation. Each approach utilizes different levers and comes with its own advantages, tradeoffs, costs, risks, a...
Organize your corporate travel faster, at lower cost. Hotailors is a next-gen AI-powered travel platform. What is Hotailors? Hotailors is a platform for organising business travels that grants access to the best real-time offers from 2.000.000+ hotels and 700+ airlines in the whole world. Thanks to our solution you can plan, book & expense business trips in less than 5 minutes. Accordingly to your travel policy, budget limits and cashless for your employees. With our reporting, int...
This sixteen (16) hour course provides an introduction to DevOps, the cultural and professional movement that stresses communication, collaboration, integration and automation in order to improve the flow of work between software developers and IT operations professionals. Improved workflows will result in an improved ability to design, develop, deploy and operate software and services faster.
Enterprises are universally struggling to understand where the new tools and methodologies of DevOps fit into their organizations, and are universally making the same mistakes. These mistakes are not unavoidable, and in fact, avoiding them gifts an organization with sustained competitive advantage, just like it did for Japanese Manufacturing Post WWII.
Andrew Keys is Co-Founder of ConsenSys Enterprise. He comes to ConsenSys Enterprise with capital markets, technology and entrepreneurial experience. Previously, he worked for UBS investment bank in equities analysis. Later, he was responsible for the creation and distribution of life settlement products to hedge funds and investment banks. After, he co-founded a revenue cycle management company where he learned about Bitcoin and eventually Ethereal. Andrew's role at ConsenSys Enterprise is a mul...
There's no doubt that blockchain technology is a powerful tool for the enterprise, but bringing it mainstream has not been without challenges. As VP of Technology at 8base, Andrei is working to make developing a blockchain application accessible to anyone. With better tools, entrepreneurs and developers can work together to quickly and effectively launch applications that integrate smart contracts and blockchain technology. This will ultimately accelerate blockchain adoption on a global scale.
DXWorldEXPO LLC announced today that Nutanix has been named "Platinum Sponsor" of CloudEXPO | DevOpsSUMMIT | DXWorldEXPO New York, which will take place November 12-13, 2018 in New York City. Nutanix makes infrastructure invisible, elevating IT to focus on the applications and services that power their business. The Nutanix Enterprise Cloud Platform blends web-scale engineering and consumer-grade design to natively converge server, storage, virtualization and networking into a resilient, softwar...