We were only just beginning to get our heads around Hadoop and Big Data in general when we find everyone is starting to talk about Big Data 2.0 – and it’s bigger, faster, and cleverer!
Hadoop, as I’m sure you know, is an open source project, and it’s available from companies like IBM, Hortonworks, Cloudera, and MapR. It provides a storage and retrieval method (HDFS – Hadoop Distributed File System) that can knock the socks off older, more expensive storage options on databases using SAN or NAS. It also means that more data can be stored. And that means not just human-keyed data, but data from the information of things (point of sales machines, sensors, cameras, etc) as well as social media. It’s an OCD sufferer’s dream come true. No need to delete (throw away) anything. But with all the data, it becomes important to find some way to ‘mine’ it – to derive information from the data that can be commercially useful. And that’s what’s happening, deeper and richer sets of results are being derived from the data that are beneficial to organizations.
With Version 2 of Hadoop, everything is faster. Data is processed at amazing speeds in-memory. The analysis is taking place at speed on terabytes of data. It also allows decisions to be made at speeds unavailable to humans. Research shows that algorithms with as many six variables out-perform human experts in most situations. This was tested on experts predicting the price of wine in future years and stock marketeers. So now, Big Data 2.0 means better decisions can be made at incredible speed.
It’s also possible for machines to learn using these techniques – such as the Google classic of having software that can identify the presence of a cat in video footage and no-one being quite sure how it is doing it.
For mainframe sites, Hadoop isn’t just some distant dream. You don’t need a room full of Linux servers to make it work – in fact that’s the clue to the solution. Much of this works very nicely on Linux on System z (or zLinux as many people still think of it). And once the data is on a mainframe, it becomes very easy to copy parts of it to a z/OS partition for more work to be done on the data. Cognos BI runs on the zLinux partition, so the first level of information extraction can be performed using that Business Intelligence tool. Software vendors are coming to market with products that run on the mainframe. BMC has extended its Control-M automated mainframe job scheduler with Control-M for Hadoop. Syncsort has Hadoop Connectivity. Compuware has extended its Application Performance Management (APM) software with Compuware APM for Big Data. And Informatica PowerExchange for Hadoop provides connectivity to the Hadoop Distributed File System (HDFS).
So what’s it like on the ground and away from the PowerPoint slides? At the moment, my experience is that really big companies – Google, Amazon, Facebook, and similar are pushing the envelope with Big Data. But it seems that many large organizations aren’t strongly embracing the new technology. Do banks, insurance companies, and airlines – the main users of mainframes – see a need for Big Data? Seemingly not – or not yet. Perhaps they are waiting for money to be spent and mistakes to be made before they adopt best practice and reap the benefits. Perhaps they are waiting for Big Data V3?
Big Data is definitely here to stay and those companies that could benefit from its adoption will gain a huge commercial advantage when they do.
Showing posts with label Cloudera. Show all posts
Showing posts with label Cloudera. Show all posts
Saturday, 1 March 2014
Sunday, 24 November 2013
Vivat mainframe
“Vivat Rex” is what the populace was meant to shout when a new king of England was crowned. It means “long live the king”. I think that we’ve been able to shout, “long live the mainframe” for a long time now, and recent announcements mean that we can continue to do so.
Mainframes, and I don’t need to tell you this, have been around for a long time now and have faced and overcome all the technical and business challenges that have been thrown at them in that time. And older mainframers can seem somewhat jaundiced when their younger colleagues get over-enthusiastic about some new technology.
We’ve looked at client-server technology and thought how similar it is to dumb terminals logging onto a mainframe. We’ve looked at cloud computing and thought how similar that is to terminals connecting to a mainframe in a different part of the world. But there’s much more to the mainframe than a simple ‘seen it, done it’ attitude. The mainframe is also able to absorb new technologies and make them its own.
We’ve looked recently at Hadoop – there are distributions from Hortonworks, Cloudera, Apache, and IBM (and many others). But you can run Big Data on your mainframe, and a number of mainframe software vendors have recently produced software that connects to Big Data from z/OS. It’s becoming integrated. So long live the mainframe with Big Data.
We’ve also, in this blog, looked at ways that BYOD – personal devices – can be used to access mainframe data, usually through browsers. And many of IBM’s younger presenters at GSE recently were talking about more Windows-like interfaces to mainframe information. Think of it – it’s like 1970 all over again – mainframes in the hands of 20-year-olds! So long live the mainframe with youthful staff and modern-looking interfaces.
We know there are other computing platforms out there, and IBM over the past few years has produced hybrid hardware that contains a mainframe and blades for running these other platforms. This summer’s zBC12 (Business Class) followed last year’s announcement of the zEC12 (Enterprise Class). And 2011 saw the z114, and 2010 gave us the z196. So long live the mainframe and its ability to embrace other platforms. (And I haven’t even mentioned how successfully you can run Linux on a mainframe.)
And thinking about mainframes embracing other technologies, CA has just announced the general availability of technology designed, they say, to help customers drive down the cost of storing data processed on IBM System z by backing up the data and archiving it to the cloud.
What that means is by using CA Cloud Storage for System z and the Riverbed Whitewater appliance, customers can back up System z storage data to Amazon Simple Storage Service (Amazon S3), a storage infrastructure designed for mission-critical and primary data storage, or to Amazon Glacier, an extremely low-cost storage service for which retrieval times of several hours are suitable. Both services are highly secure and scalable and designed to be durable. In addition, disaster recovery readiness is improved and AWS cloud storage is accessed without changing the existing back-up infrastructure.
So, yet again, we can say, long live the mainframe for the way it’s embracing cloud computing.
As a side note: Amazon has Amazon Elastic MapReduce (EMR), which uses Hadoop to provide Web services.
IBM has taken over StoredIQ, Star Analytics, and The Now Factory for Big Data Analytics or Business Analytics. And it took over SoftLayer Technologies for its cloud computing infrastructure. It’s making sure it has its hands on the tools and the people who are developing these newer technologies.
My conclusion is that there are new problems that need to be solved. And there are new technologies available to solve them. But so often those exciting new things are very similar to things that we mainframers have dealt with before. And where they seem different, mainframe environments are able to work with them and bring them into the fold.
There’s really no danger that mainframes are going away anytime soon. So, we’re very safe in saying, “vivat mainframe”.
On a completely different topic...
Please complete the mainframe users’ survey at www.arcati.com/usersurvey14. And if you’re a vendor, get your free entry in the Arcati Mainframe Yearbook 2014 by completing the form at www.arcati.com/vendorentry.
Mainframes, and I don’t need to tell you this, have been around for a long time now and have faced and overcome all the technical and business challenges that have been thrown at them in that time. And older mainframers can seem somewhat jaundiced when their younger colleagues get over-enthusiastic about some new technology.
We’ve looked at client-server technology and thought how similar it is to dumb terminals logging onto a mainframe. We’ve looked at cloud computing and thought how similar that is to terminals connecting to a mainframe in a different part of the world. But there’s much more to the mainframe than a simple ‘seen it, done it’ attitude. The mainframe is also able to absorb new technologies and make them its own.
We’ve looked recently at Hadoop – there are distributions from Hortonworks, Cloudera, Apache, and IBM (and many others). But you can run Big Data on your mainframe, and a number of mainframe software vendors have recently produced software that connects to Big Data from z/OS. It’s becoming integrated. So long live the mainframe with Big Data.
We’ve also, in this blog, looked at ways that BYOD – personal devices – can be used to access mainframe data, usually through browsers. And many of IBM’s younger presenters at GSE recently were talking about more Windows-like interfaces to mainframe information. Think of it – it’s like 1970 all over again – mainframes in the hands of 20-year-olds! So long live the mainframe with youthful staff and modern-looking interfaces.
We know there are other computing platforms out there, and IBM over the past few years has produced hybrid hardware that contains a mainframe and blades for running these other platforms. This summer’s zBC12 (Business Class) followed last year’s announcement of the zEC12 (Enterprise Class). And 2011 saw the z114, and 2010 gave us the z196. So long live the mainframe and its ability to embrace other platforms. (And I haven’t even mentioned how successfully you can run Linux on a mainframe.)
And thinking about mainframes embracing other technologies, CA has just announced the general availability of technology designed, they say, to help customers drive down the cost of storing data processed on IBM System z by backing up the data and archiving it to the cloud.
What that means is by using CA Cloud Storage for System z and the Riverbed Whitewater appliance, customers can back up System z storage data to Amazon Simple Storage Service (Amazon S3), a storage infrastructure designed for mission-critical and primary data storage, or to Amazon Glacier, an extremely low-cost storage service for which retrieval times of several hours are suitable. Both services are highly secure and scalable and designed to be durable. In addition, disaster recovery readiness is improved and AWS cloud storage is accessed without changing the existing back-up infrastructure.
So, yet again, we can say, long live the mainframe for the way it’s embracing cloud computing.
As a side note: Amazon has Amazon Elastic MapReduce (EMR), which uses Hadoop to provide Web services.
IBM has taken over StoredIQ, Star Analytics, and The Now Factory for Big Data Analytics or Business Analytics. And it took over SoftLayer Technologies for its cloud computing infrastructure. It’s making sure it has its hands on the tools and the people who are developing these newer technologies.
My conclusion is that there are new problems that need to be solved. And there are new technologies available to solve them. But so often those exciting new things are very similar to things that we mainframers have dealt with before. And where they seem different, mainframe environments are able to work with them and bring them into the fold.
There’s really no danger that mainframes are going away anytime soon. So, we’re very safe in saying, “vivat mainframe”.
On a completely different topic...
Please complete the mainframe users’ survey at www.arcati.com/usersurvey14. And if you’re a vendor, get your free entry in the Arcati Mainframe Yearbook 2014 by completing the form at www.arcati.com/vendorentry.
Labels:
Amazon,
Apache,
big data,
blog,
CA Cloud Storage for System z,
cloud,
Cloudera,
Eddolls,
Hadoop,
Hortonworks,
IBM,
mainframe,
Riverbed Whitewater appliance,
S3
Sunday, 2 June 2013
Big data – where are we?
At first, people would enter information into their computers, then print it off if they wanted to share the data. Then we had networks and people could electronically share data – and then others could add to it. Pretty much all the data – even in the largest IMS database – had been entered by people or calculated from data entered by people.
But more recently, things have changed. Information stored on computers has come from other sources, for example card readers, CCTV cameras, traffic flow sensors, etc, etc. Almost any device can be given an IP address, connected to a network, and used as a source of data. All these ‘things’, that can and are being connected, has led to the use of the phrase: ‘the Internet of things’. Perhaps not the most precise description, but it indicates that the Internet is being used as a way of getting information from devices – rather than waiting for a human to type in the data.
The other development that we’re all familiar with is the growth in cloud computing. What that means is devices are connected to a nebulous source of storage and processing power. Mainframers, who have been around the block a few times, feel quite happy with this model of dumb terminals connected to some giant processing device that is some distance away and not necessarily visible to the users of the dumb terminals. This is what mainframe computing was like (and still is for some users!). Other computer professionals will recognize this as another version of the client/server model that was once so fashionable.
By having so many sources of data input, you have security and storage issues, but, perhaps more importantly, you have issues about what to do with the data. It’s almost like a person with OCD hoarding old newspaper that they never look at but can’t throw away. What can you do with these vast amounts of data?
The answer is Hadoop. According to the Web site at http://hadoop.apache.org/: “The Apache Hadoop software library is a framework that allows for the distributed processing of large data sets across clusters of computers using simple programming models. It is designed to scale up from single servers to thousands of machines, each offering local computation and storage. Rather than rely on hardware to deliver high-availability, the library itself is designed to detect and handle failures at the application layer, so delivering a highly-available service on top of a cluster of computers, each of which may be prone to failures.”
So which companies are experienced with Hadoop? Cloudera was probably the best known in the field up until recently. Other companies you may not have heard of are MapR and Hortonworks. Companies you will be familiar with are EMC and VMware who have spun off a company called Pivotal. And there’s Intel, and there’s IBM.
Let’s have a quick look at what’s out there. Apache Hive was developed by Facebook, but is now Open Source. Dremel (from Google) is published, but not yet available. Apache Drill is based on Dremel, but is still in the incubation stage. Cloudera’s Impala was inspired by Dremel. IBM’s offering is Big SQL. Hive is a data warehouse infrastructure built on top of Hadoop. It converts queries into MapReduce jobs. Impala’s SQL query system for Hadoop is Open Source. It uses C++ rather than Java. It doesn’t use MapReduce. Impala only works with Cloudera’s Distribution of Hadoop (CDH).
The Apache Thrift software framework, for scalable cross-language services development, combines a software stack with a code generation engine to build services that work efficiently and seamlessly between C++, Java, Python, PHP, Ruby, Erlang, Perl, Haskell, C#, Cocoa, JavaScript, Node.js, Smalltalk, OCaml, and Delphi and other languages.
IBM’s Big SQL is a currently a technology preview. It supports SQL, and JDBC and ODBC client drivers. IBM’s distribution of Hadoop is called BigInsights. Big SQL is similar to Hive and they can cross query. Point query is used for small queries rather than MapReduce. It supports more datatypes than Hive.
So, you can see that there’s lot’s to learn about Hadoop, and I’m sure we’ll be hearing a lot more about BigInsights and Big SQL. My advice is, if you’re looking for a career path, companies are going to need experienced Hadoop people – so get some!
But more recently, things have changed. Information stored on computers has come from other sources, for example card readers, CCTV cameras, traffic flow sensors, etc, etc. Almost any device can be given an IP address, connected to a network, and used as a source of data. All these ‘things’, that can and are being connected, has led to the use of the phrase: ‘the Internet of things’. Perhaps not the most precise description, but it indicates that the Internet is being used as a way of getting information from devices – rather than waiting for a human to type in the data.
The other development that we’re all familiar with is the growth in cloud computing. What that means is devices are connected to a nebulous source of storage and processing power. Mainframers, who have been around the block a few times, feel quite happy with this model of dumb terminals connected to some giant processing device that is some distance away and not necessarily visible to the users of the dumb terminals. This is what mainframe computing was like (and still is for some users!). Other computer professionals will recognize this as another version of the client/server model that was once so fashionable.
By having so many sources of data input, you have security and storage issues, but, perhaps more importantly, you have issues about what to do with the data. It’s almost like a person with OCD hoarding old newspaper that they never look at but can’t throw away. What can you do with these vast amounts of data?
The answer is Hadoop. According to the Web site at http://hadoop.apache.org/: “The Apache Hadoop software library is a framework that allows for the distributed processing of large data sets across clusters of computers using simple programming models. It is designed to scale up from single servers to thousands of machines, each offering local computation and storage. Rather than rely on hardware to deliver high-availability, the library itself is designed to detect and handle failures at the application layer, so delivering a highly-available service on top of a cluster of computers, each of which may be prone to failures.”
So which companies are experienced with Hadoop? Cloudera was probably the best known in the field up until recently. Other companies you may not have heard of are MapR and Hortonworks. Companies you will be familiar with are EMC and VMware who have spun off a company called Pivotal. And there’s Intel, and there’s IBM.
Let’s have a quick look at what’s out there. Apache Hive was developed by Facebook, but is now Open Source. Dremel (from Google) is published, but not yet available. Apache Drill is based on Dremel, but is still in the incubation stage. Cloudera’s Impala was inspired by Dremel. IBM’s offering is Big SQL. Hive is a data warehouse infrastructure built on top of Hadoop. It converts queries into MapReduce jobs. Impala’s SQL query system for Hadoop is Open Source. It uses C++ rather than Java. It doesn’t use MapReduce. Impala only works with Cloudera’s Distribution of Hadoop (CDH).
The Apache Thrift software framework, for scalable cross-language services development, combines a software stack with a code generation engine to build services that work efficiently and seamlessly between C++, Java, Python, PHP, Ruby, Erlang, Perl, Haskell, C#, Cocoa, JavaScript, Node.js, Smalltalk, OCaml, and Delphi and other languages.
IBM’s Big SQL is a currently a technology preview. It supports SQL, and JDBC and ODBC client drivers. IBM’s distribution of Hadoop is called BigInsights. Big SQL is similar to Hive and they can cross query. Point query is used for small queries rather than MapReduce. It supports more datatypes than Hive.
So, you can see that there’s lot’s to learn about Hadoop, and I’m sure we’ll be hearing a lot more about BigInsights and Big SQL. My advice is, if you’re looking for a career path, companies are going to need experienced Hadoop people – so get some!
Subscribe to:
Posts (Atom)