Showing posts with label Hadoop. Show all posts
Showing posts with label Hadoop. Show all posts

Sunday, 27 April 2014

Tell me about NoSQL

NoSQL seems to be the buzzword of choice at the moment for people who want the flexibility to build and frequently alter their databases. But there are plenty of people who still aren’t quite sure what a NoSQL database is and why they should want to use it. So let’s take a brief overview of NoSQL.

The term, NoSQL, first saw the light of day in 1998 when Carlo Strozzi used it as the name of his lightweight, open-source, relational database because it didn’t expose the standard SQL interface. But the term gained its modern usage in 2009 when it was used as a generic label for non-relational, distributed, data stores. So, it refers to a whole family of databases, rather than a single type of database.

Developers like NoSQL because they can store and retrieve data without being locked into the tabular relationships used in relational databases. It makes scaling easier and they provide superior performance. They can store large volumes of structured, semi-structured, and unstructured data. They can handle agile sprints, quick iteration, and frequent code pushes. They use object-oriented programming that is easy to use and flexible. And they use efficient scale-out architecture instead of expensive monolithic architecture.

But, on the down side, NoSQL lacks ACID (Atomicity, Consistency, Isolation, Durability) transaction support. Atomicity means that each transaction is ‘all or nothing’, ie if one part of the transaction fails, the entire transaction fails, and the database state is left unchanged. Consistency ensures that any transaction brings the database from one valid state to another. Isolation means that the concurrent execution of transactions results in a system state that would be obtained if transactions were executed sequentially. Durability means that once a transaction has been committed, it will remain so, even in the event of power loss, crashes, or errors. And that’s the kind of reliability you want in a business-critical database.

NoSQL databases are typically used in Big Data and real-time Web applications. The different NoSQL database technologies were developed because of the increase in the volume of data that people needed to store, the frequency the data is accessed, and increased performance and processing needs.

There are estimated to be over 150 open source databases available. And there are many different types of NoSQL database, including some that allow the use of SQL-like languages – these are sometimes referred to as ‘Not only SQL’ databases. Classify NoSQL databases is quite a challenge, but they can be grouped, by the features they offer, into column, document, key-value, and graph types. Alternatively, they can be classified by data model into KV Cache, KV Store, KV Store - Eventually consistent, Data-structures server, KV Store – Ordered, Tuple Store, Object Database, Document Store, and Wide Columnar Store.

The good news for DB2 users is that IBM has provided a new API that supports multiple calls and a NoSQL software solution stack that ships with DB2. It’s free with DB2 on distributed platforms and with DB2 Connect. DB2 also offers a second type of NoSQL-like database – the XML data store. This can store the growing volume of Web-based data.

Rocket Software has a way of using MongoDB (an example of a NoSQL database) on a mainframe. Rocket can provide access to any System z database using any MongoDB client driver. DB2 supports MongoDB.

IBM recently announced Zdoop, Hadoop database software for Linux from Veristorm on System z mainframes, stating: “This will help clients to avoid staging and offloading mainframe data to maintain existing security and governance controls”.

Other NoSQL databases that you might want to look out for include Cassandra, CouchBase, Redis, and Riak.

Clearly, with the growth in Big Data, we’ll be hearing a lot more about NoSQL databases and how they can be integrated into mainframe technology. There are lots of them out there and they can be quite different from each other in terms of their features and data models used.

Saturday, 1 March 2014

Big Data 2.0

We were only just beginning to get our heads around Hadoop and Big Data in general when we find everyone is starting to talk about Big Data 2.0 – and it’s bigger, faster, and cleverer!

Hadoop, as I’m sure you know, is an open source project, and it’s available from companies like IBM, Hortonworks, Cloudera, and MapR. It provides a storage and retrieval method (HDFS – Hadoop Distributed File System) that can knock the socks off older, more expensive storage options on databases using SAN or NAS. It also means that more data can be stored. And that means not just human-keyed data, but data from the information of things (point of sales machines, sensors, cameras, etc) as well as social media. It’s an OCD sufferer’s dream come true. No need to delete (throw away) anything. But with all the data, it becomes important to find some way to ‘mine’ it – to derive information from the data that can be commercially useful. And that’s what’s happening, deeper and richer sets of results are being derived from the data that are beneficial to organizations.

With Version 2 of Hadoop, everything is faster. Data is processed at amazing speeds in-memory. The analysis is taking place at speed on terabytes of data. It also allows decisions to be made at speeds unavailable to humans. Research shows that algorithms with as many six variables out-perform human experts in most situations. This was tested on experts predicting the price of wine in future years and stock marketeers. So now, Big Data 2.0 means better decisions can be made at incredible speed.

It’s also possible for machines to learn using these techniques – such as the Google classic of having software that can identify the presence of a cat in video footage and no-one being quite sure how it is doing it.

For mainframe sites, Hadoop isn’t just some distant dream. You don’t need a room full of Linux servers to make it work – in fact that’s the clue to the solution. Much of this works very nicely on Linux on System z (or zLinux as many people still think of it). And once the data is on a mainframe, it becomes very easy to copy parts of it to a z/OS partition for more work to be done on the data. Cognos BI runs on the zLinux partition, so the first level of information extraction can be performed using that Business Intelligence tool. Software vendors are coming to market with products that run on the mainframe. BMC has extended its Control-M automated mainframe job scheduler with Control-M for Hadoop. Syncsort has Hadoop Connectivity. Compuware has extended its Application Performance Management (APM) software with Compuware APM for Big Data. And Informatica PowerExchange for Hadoop provides connectivity to the Hadoop Distributed File System (HDFS).

So what’s it like on the ground and away from the PowerPoint slides? At the moment, my experience is that really big companies – Google, Amazon, Facebook, and similar are pushing the envelope with Big Data. But it seems that many large organizations aren’t strongly embracing the new technology. Do banks, insurance companies, and airlines – the main users of mainframes – see a need for Big Data? Seemingly not – or not yet. Perhaps they are waiting for money to be spent and mistakes to be made before they adopt best practice and reap the benefits. Perhaps they are waiting for Big Data V3?

Big Data is definitely here to stay and those companies that could benefit from its adoption will gain a huge commercial advantage when they do.

Sunday, 26 January 2014

The Arcati Mainframe Yearbook - user survey findings

The Arcati Mainframe Yearbook 2014 is now available for download from http://www.arcati.com/newyearbook14 – and it’s FREE. Each new Yearbook is always greeted with enthusiasm by mainframers everywhere because it is such a unique source of information. And each year, many people find the results of the user survey especially interesting.

The results came from the 100 respondents who completed the survey on the Arcati Web site between 1 November and 6 December 2013. 51% were from North America, 33% were from Europe with the remainder from the rest of the world.

Half of the respondents worked in companies with upwards of 10,000 employees worldwide. Below that, with 24 percent of respondents, were staff sizes of 1001-5000, 10 percent with staff sizes of 0-200, nine percent with staff sizes of 201 to 1000, and only seven percent with staff sizes of 5001 to 10,000. In terms of MIPS, 28 percent had 1000-10,000 MIPS, down again from last year’s figure of 36 percent. 16 percent had under 500 MIPS, only 15 percent had 500-1000 MIPS, 13 percent had 10,000 to 25,000 MIPS, and 17 percent had over 25,000 MIPS installed.

Looking at MIPS growth produced some interesting results. 71 percent of sites of mainframe installations are experiencing some growth, with three sites claiming growth in the region of 26-50 percent. Only eight percent of sites are reporting a decline in mainframe capacity growth. 11 percent of sites are not expecting any kind of change in their MIPS this year. Small sites (32 percent) are most likely to have seen some kind of decline or to have stayed the same, and yet, in complete contrast, they were more likely to see growth in the 26-50 percent range. While some larger sites (above 10,000 MIPS) did report a decline or no growth, the majority were anticipating some kind of growth possibly up to 50 percent per year. Mid-range respondents were typically expecting some kind of growth (89 percent of sites). It is a confusing picture with nearly a third of small sites, 10 percent of medium sites, and 17 percent of larger sites showing no growth or a decline. Perhaps sites have been holding off on growth until the global economic climate brightens up.

The survey looked at whether sites currently used their mainframe for cloud computing. Only seven percent of respondents said they did. The survey also asked whether respondents were planning to adopt cloud computing as a strategy. 50 percent said they weren’t at present. 16 percent thought some mainframe applications would be cloud enabled in the future. And 15 percent  claimed that some of their applications are using the cloud model.

There’s been a huge growth in the use of social media in recent years, and the survey wondered whether those people “using their dad’s technology” found social media (Facebook, Twitter, Youtube, etc) useful for their work on the mainframe. 18 percent said that they did, with 13 percent not sure, and the rest not using it at all. With IBM having Facebook pages dedicated to IMS, CICS, and DB2, it seems a shame if they’re not being used.

With the growth in number of software products that allow users to monitor the mainframe from a browser on a tablet/iPad or smartphone, the survey looked at whether mainframers were using these devices to monitor or control their mainframe. Only nine percent said that they were.

Another hot topic through 2013 has been Big Data and all the things associated with that (such as Hadoop). The survey asked whether sites had any plans to use Big Data. Just two percent of sites said that they were already using Big Data, with a further 12 percent planning to do so.

The survey also asked about BYOD (Bring Your Own Device). It wanted to know how important sites thought it was to make mainframe data available to other platforms. 78 percent of sites said that it was very important to the way they work at the moment. Three percent are in the planning stage, and nine percent expect to do some work on this in the future. When it comes to how important is the idea of people using their own devices (BYOD) to access mainframes, 15 percent of sites said it was very important to the way they work now – but 47 percent said it wasn’t important.

Anyway, full details of the responses to many other questions can be found in the user survey section of the Yearbook. It’s well worth a read.

The Yearbook can only be free because some organizations have been prepared to sponsor it or advertise in it. This year’s sponsors were: Software Diversified Services (SDS), Software AG, zIT Consulting, and CA Technologies.

Sunday, 24 November 2013

Vivat mainframe

“Vivat Rex” is what the populace was meant to shout when a new king of England was crowned. It means “long live the king”. I think that we’ve been able to shout, “long live the mainframe” for a long time now, and recent announcements mean that we can continue to do so.

Mainframes, and I don’t need to tell you this, have been around for a long time now and have faced and overcome all the technical and business challenges that have been thrown at them in that time. And older mainframers can seem somewhat jaundiced when their younger colleagues get over-enthusiastic about some new technology.

We’ve looked at client-server technology and thought how similar it is to dumb terminals logging onto a mainframe. We’ve looked at cloud computing and thought how similar that is to terminals connecting to a mainframe in a different part of the world. But there’s much more to the mainframe than a simple ‘seen it, done it’ attitude. The mainframe is also able to absorb new technologies and make them its own.

We’ve looked recently at Hadoop – there are distributions from Hortonworks, Cloudera, Apache, and IBM (and many others). But you can run Big Data on your mainframe, and a number of mainframe software vendors have recently produced software that connects to Big Data from z/OS. It’s becoming integrated. So long live the mainframe with Big Data.

We’ve also, in this blog, looked at ways that BYOD – personal devices – can be used to access mainframe data, usually through browsers. And many of IBM’s younger presenters at GSE recently were talking about more Windows-like interfaces to mainframe information. Think of it – it’s like 1970 all over again – mainframes in the hands of 20-year-olds! So long live the mainframe with youthful staff and modern-looking interfaces.

We know there are other computing platforms out there, and IBM over the past few years has produced hybrid hardware that contains a mainframe and blades for running these other platforms. This summer’s zBC12 (Business Class) followed last year’s announcement of the zEC12 (Enterprise Class). And 2011 saw the z114, and 2010 gave us the z196. So long live the mainframe and its ability to embrace other platforms. (And I haven’t even mentioned how successfully you can run Linux on a mainframe.)

And thinking about mainframes embracing other technologies, CA has just announced the general availability of technology designed, they say, to help customers drive down the cost of storing data processed on IBM System z by backing up the data and archiving it to the cloud.

What that means is by using CA Cloud Storage for System z and the Riverbed Whitewater appliance, customers can back up System z storage data to Amazon Simple Storage Service (Amazon S3), a storage infrastructure designed for mission-critical and primary data storage, or to Amazon Glacier, an extremely low-cost storage service for which retrieval times of several hours are suitable. Both services are highly secure and scalable and designed to be durable. In addition, disaster recovery readiness is improved and AWS cloud storage is accessed without changing the existing back-up infrastructure.
So, yet again, we can say, long live the mainframe for the way it’s embracing cloud computing.

As a side note: Amazon has Amazon Elastic MapReduce (EMR), which uses Hadoop to provide Web services.

IBM has taken over StoredIQ, Star Analytics, and The Now Factory for Big Data Analytics or Business Analytics. And it took over SoftLayer Technologies for its cloud computing infrastructure. It’s making sure it has its hands on the tools and the people who are developing these newer technologies.

My conclusion is that there are new problems that need to be solved. And there are new technologies available to solve them. But so often those exciting new things are very similar to things that we mainframers have dealt with before. And where they seem different, mainframe environments are able to work with them and bring them into the fold.

There’s really no danger that mainframes are going away anytime soon. So, we’re very safe in saying, “vivat mainframe”.

On a completely different topic...
Please complete the mainframe users’ survey at www.arcati.com/usersurvey14. And if you’re a vendor, get your free entry in the Arcati Mainframe Yearbook 2014 by completing the form at www.arcati.com/vendorentry.

Sunday, 7 July 2013

IBM’s approach to Big Data

IBM has taken lots of the open source Big Data technologies – like Hadoop, MapReduce, HBase – and added its own technology – like Big Sheets, DB2, DataStage – to create something hugely more powerful.

IBM’s InfoSphere BigInsights builds on open source Hadoop capabilities for enterprise class deployments. The enterprise-level capabilities can be grouped together as: visualization and exploration, development tools, advanced engines, connectors, workload optimization, and administration and security.

IBM claims the business benefits are: quicker time-to-value because of IBM’s technology and support, reduced operational risk, enhanced business knowledge with a flexible analytical platform, and it leverages and complements existing software.

In terms of administration and security, the Web console can start and stop services, run and monitor jobs (applications), explore and modify the file system, and built-in apps make it easy to do common tasks.

The connectors link to databases like DB2, Netezza, Oracle, Teradata. And there’s integration with: InfoSphere Data Stage (data collection and integration), InfoSphere Streams (real-time streams processing), InfoSphere Guardium (security and monitoring), Cognos Business Intelligence (Business Intelligence capabilities), and IBM Platform Computing (cluster/grid infrastructure and management), and more. Big SQL is coming with BigInsights V2.1. This will provide SQL access to data stored in BigInsights through JDBC/ODBC and use rich standard SQL to leverage Map/Reduce parallelism or achieve low-latency.

Advanced engines include an advanced text analytics engine that can automatically identify and understand key information in text. Text Analytics is really useful because most of the world’s data is in unstructured or semi-structured text; social media is full of discussions about products and services; internal information in organizations is locked in blobs, description fields, and sometimes even discarded. It’s been suggested that over 80% of stored information is unstructured – such as e-medical records, hospital reports, case files, police records, emergency calls, tech notes, call logs, online media, insurance claims, Twitter, Facebook, blogs, and forums.

In terms of development tools, there is an Eclipse-based development environment for building and deploying applications. There are developer tools and a set of analytic extractors for fast adoption that reduce coding and debugging time by up to 30% (IBM claims). There are also plug-ins for text analytics, MapReduce programming, Jaql development, Hive query, etc.

Visualization and exploration has Big Sheets, providing Web-based analysis and visualization for users with a familiar spreadsheet-like interface that can define and manage long-running data collection jobs.

Meanwhile, Microsoft has identified Hadoop users as a useful market to get into. Speaking recently at the Hadoop summit, Quentin Clark, corporate VP of data platforms said: “We believe Hadoop is the cornerstone of a sea change coming to all businesses”.

Microsoft is integrating Hadoop with its products and services. And, Clark says that Microsoft intends to stick to the principles of open source by contributing to the Hadoop project, rather than simply using it and adding its own stuff. Hortonworks recently announced management packs for Microsoft System Center Operations Manager and Microsoft System Center Virtual Machine Manager – both products for administering the Hortonworks Data Platform (HDP) distribution.

Apparently Microsoft is positioning itself as a big data player with a powerful set of Business Intelligence (BI) tools. Data Explorer for Excel 2013 is a self-service BI add-in allowing users to import data from a variety of sources, including Hadoop. SQL Server 2012 Parallel Data Warehouse (PDW) is a massively parallel processing data warehousing appliance designed for Hadoop integration. Microsoft is also trying to bring Hadoop into the cloud using Windows Azure.

Businesses can’t ignore Hadoop, and the fact that major software vendors are getting behind it means it’s not going to be some flash-in-the-pan idea. Certainly, I can imagine major organizations looking to get a huge business advantage by embracing the technology now – to be ahead of their competitors. Smaller organizations will probably take a few years before they see a business case for it. By then the IBM products (and Microsoft’s) will be very mature and eminently suitable.

Sunday, 2 June 2013

Big data – where are we?

At first, people would enter information into their computers, then print it off if they wanted to share the data. Then we had networks and people could electronically share data – and then others could add to it. Pretty much all the data – even in the largest IMS database – had been entered by people or calculated from data entered by people.

But more recently, things have changed. Information stored on computers has come from other sources, for example card readers, CCTV cameras, traffic flow sensors, etc, etc. Almost any device can be given an IP address, connected to a network, and used as a source of data. All these ‘things’, that can and are being connected, has led to the use of the phrase: ‘the Internet of things’. Perhaps not the most precise description, but it indicates that the Internet is being used as a way of getting information from devices – rather than waiting for a human to type in the data.

The other development that we’re all familiar with is the growth in cloud computing. What that means is devices are connected to a nebulous source of storage and processing power. Mainframers, who have been around the block a few times, feel quite happy with this model of dumb terminals connected to some giant processing device that is some distance away and not necessarily visible to the users of the dumb terminals. This is what mainframe computing was like (and still is for some users!). Other computer professionals will recognize this as another version of the client/server model that was once so fashionable.

By having so many sources of data input, you have security and storage issues, but, perhaps more importantly, you have issues about what to do with the data. It’s almost like a person with OCD hoarding old newspaper that they never look at but can’t throw away. What can you do with these vast amounts of data?

The answer is Hadoop. According to the Web site at http://hadoop.apache.org/: “The Apache Hadoop software library is a framework that allows for the distributed processing of large data sets across clusters of computers using simple programming models. It is designed to scale up from single servers to thousands of machines, each offering local computation and storage. Rather than rely on hardware to deliver high-availability, the library itself is designed to detect and handle failures at the application layer, so delivering a highly-available service on top of a cluster of computers, each of which may be prone to failures.”

So which companies are experienced with Hadoop? Cloudera was probably the best known in the field up until recently. Other companies you may not have heard of are MapR and Hortonworks. Companies you will be familiar with are EMC and VMware who have spun off a company called Pivotal. And there’s Intel, and there’s IBM.

Let’s have a quick look at what’s out there. Apache Hive was developed by Facebook, but is now Open Source. Dremel (from Google) is published, but not yet available. Apache Drill is based on Dremel, but is still in the incubation stage. Cloudera’s Impala was inspired by Dremel. IBM’s offering is Big SQL. Hive is a data warehouse infrastructure built on top of Hadoop. It converts queries into MapReduce jobs. Impala’s SQL query system for Hadoop is Open Source. It uses C++ rather than Java. It doesn’t use MapReduce. Impala only works with Cloudera’s Distribution of Hadoop (CDH).

The Apache Thrift software framework, for scalable cross-language services development, combines a software stack with a code generation engine to build services that work efficiently and seamlessly between C++, Java, Python, PHP, Ruby, Erlang, Perl, Haskell, C#, Cocoa, JavaScript, Node.js, Smalltalk, OCaml, and Delphi and other languages.

IBM’s Big SQL is a currently a technology preview. It supports SQL, and JDBC and ODBC client drivers. IBM’s distribution of Hadoop is called BigInsights. Big SQL is similar to Hive and they can cross query. Point query is used for small queries rather than MapReduce. It supports more datatypes than Hive.

So, you can see that there’s lot’s to learn about Hadoop, and I’m sure we’ll be hearing a lot more about BigInsights and Big SQL. My advice is, if you’re looking for a career path, companies are going to need experienced Hadoop people – so get some!

Saturday, 6 April 2013

Big SQL

Suppose you wanted to access ‘big data’ stored in HDFS or HBase. What would you do? Well, for many people, the first step is to find out what we’re talking about. So, let’s start with big data – it’s data that’s so large and complex that it’s difficult to process using standard and familiar database management tools or applications.

According to Wikipedia, there are issues around data capture, curation, storage, search, sharing, analysis, and visualization. You’re probably thinking: why not go back to using smaller and manageable data? It seems that people want access to larger and larger amounts of data because additional information can be gained from it – allowing people to “spot business trends, determine quality of research, prevent diseases, link legal citations, combat crime, and determine real-time roadway traffic conditions”.

Now that’s clear, what are HDFS and HBase? HDFS stands for Hadoop Distributed File System. It’s a distributed, scalable, and portable file system written in Java for the Hadoop framework. HDFS stores large files across multiple machines, and replicates the data across multiple hosts. HBase is an open source, non-relational, distributed database and is also written in Java. It was developed as part of Apache Software Foundation’s Apache Hadoop project and runs on top of HDFS (Hadoop Distributed File System), providing a fault-tolerant way of storing large quantities of data.

Each node in a Hadoop instance typically has a single namenode; a cluster of datanodes form the HDFS cluster. So what’s needed is some way to access that cluster. At the moment, the choices are basically Hive, Impala, and Big SQL.

Again, a search on Wikipedia informs me that “Hive supports analysis of large datasets stored in Hadoop-compatible file systems such as Amazon S3 filesystem. It provides an SQL-like language called HiveQL while maintaining full support for map/reduce. To accelerate queries, it provides indexes, including bitmap indexes. By default, Hive stores metadata in an embedded Apache Derby database, and other client/server databases like MySQL can optionally be used. Currently, there are three file formats supported in Hive, which are TEXTFILE, SEQUENCEFILE, and RCFILE”

The Cloudera Impala project allows users to query data, whether stored in HDFS or HBase – including SELECT, JOIN, and aggregate functions – in real time. Furthermore, it uses the same metadata, SQL syntax (Hive SQL), ODBC driver, and user interface (Hue Beeswax) as Apache Hive. To avoid latency, Impala circumvents MapReduce to directly access the data through a specialized distributed query engine.

When you look up information about these things, names like Apache, Cloudera, Amazon, Facebook, Google crop up, but not IBM. You might think that’s a bit strange. Wouldn’t IBM be the organization you’d expect to have experience of big data? I mean just think of those massive IMS databases. So, why haven’t I mentioned IBM? The answer is because I haven’t got to Big SQL yet.

IBM claims that Big SQL provides robust SQL support for the Hadoop ecosystem:

  •  it has a scalable architecture;
  • it supports SQL and data types available in SQL '92, plus it has some additional capabilities;
  • it supports JDBC and ODBC client drivers;
  • it has efficient handling of ‘point queries’;
  • there are a wide variety of data sources and file formats for HDFS and HBase that it supports;
  • And, although it isn’t open source, it does interoperate well with the open source ecosystem within Hadoop.

The really interesting thing about this is that all the information is available in one place – Big Data University (http://bigdatauniversity.com). I’m looking forward to taking the course. Big data isn’t going away any time soon.

Saturday, 18 August 2012

Why is everyone talking about Hadoop?

Hadoop is an Apache project, which means it’s open source software, and it’s written in Java. What it does is support data-intensive distributed applications. It comes from work Google were doing and allows applications to use thousands of independent computers and petabytes of data.

Yahoo has been a big contributor to the project. The Yahoo Search Webmap is a Hadoop application that is used in every Yahoo search. Facebook claims to have the largest Hadoop cluster in the world. Other users include Amazon, eBay, LinkedIn, and Twitter. But now, there’s talk of IBM taking more than a passing interest.

According to IBM: “Apache Hadoop has two main subprojects:
  • MapReduce – The framework that understands and assigns work to the nodes in a cluster.
  • HDFS – A file system that spans all the nodes in a Hadoop cluster for data storage. It links together the file systems on many local nodes to make them into one big file system. HDFS assumes nodes will fail, so it achieves reliability by replicating data across multiple nodes.”

It goes on to say: “Hadoop changes the economics and the dynamics of large-scale computing. Its impact can be boiled down to four salient characteristics. Hadoop enables a computing solution that is:
  • Scalable – New nodes can be added as needed, and added without needing to change data formats, how data is loaded, how jobs are written, or the applications on top.
  • Cost effective – Hadoop brings massively parallel computing to commodity servers. The result is a sizeable decrease in the cost per terabyte of storage, which in turn makes it affordable to model all your data.
  • Flexible – Hadoop is schema-less, and can absorb any type of data, structured or not, from any number of sources. Data from multiple sources can be joined and aggregated in arbitrary ways enabling deeper analyses than any one system can provide.
  • Fault tolerant – When you lose a node, the system redirects work to another location of the data and continues processing without missing a beat.”

According to Alan Radding writing in IBM Systems Magazine (http://www.ibmsystemsmag.com/mainframe/trends/whatsnew/hadoop_mainframe/) IBM “is taking a federated approach to the big data challenge by blending traditional data management technologies with what it sees as complementary new technologies, like Hadoop, that address speed and flexibility, and are ideal for data exploration, discovery and unstructured analysis.”

Hadoop could run on any mainframe already running Java or Linux. Radding lists tools to make life easier like:
  • SQOOP – imports data from relational databases into Hadoop.
  • Hive – enables data to be queried using an SQL-like language called HiveQL.
  • Apache Pig – a high-level platform for creating the MapReduce programs used with Hadoop.

There’s also ZooKeeper, which provides a centralized infrastructure and services that enable synchronization across a cluster.

Harry Battan, data serving manager for System z, suggests that 2,000 instances of Hadoop could run on Linux on the System z, which would make a fairly large Hadoop configuration.

Hadoop still needs to be certified for mainframe use, but sites with newer hybrid machines (z114 or z196) could have Hadoop today by putting it on their x86 blades, for which Hadoop is already certified, and it could then process data from DB2 on the mainframe. But you can see why customers might be looking to get it on their mainframes because it gives them a way to get more information out of the masses of data they already possess. And data analysis is often seen as the key to continuing business success for larger organizations.