Sunday, 22 August 2021

Mainframe resilience


Have you ever been part of a Business Continuity Plan test? If you have, then you know that you tend to end up in a hotel, or some other building, with a variety of other people from an organization and ‘war game’ what would happen in the event of various scenarios. Often, an external company will be invited in to host the sessions and be on the other end of the phone when someone is trying to deal with the press. The day can often be quite fun, sometimes illuminating, and lunch is usually very good!

The big problem is that often, what can be resolved in the meeting room in a couple of minutes takes much much longer in real life. In many scenarios, a building has been burgled, or is full of terrorists, but the mainframe and the other servers are still working. If the communications line from one site have been cut, most people – as we’ve seen over the past year or so – can work from home. The organization is generally able to continue in business so long as the mainframe is still working. But what happens if it isn’t?

I can remember many years ago, at one site I worked at, putting a tick in the box for backups for a particular application. However, as all the operators knew, the backup tapes were 7-track tapes, and the last 7-track tape drive had been removed some months beforehand. There was no way that anything could be restored. I can also remember driving backup tapes to an offsite backup site at a company on the other side of town. If there had been a disaster out of hours, can you imagine how long it would have taken to get those tapes and restore the data?

Clearly, backup strategies have improved hugely since those days back in the early 1980s. Even so, a lot of emphasis is still being put on the backing up of data, and, all too often, not enough emphasis is put on restoring the data. I’m talking about mainframe resiliency.

Mainframe resiliency is the ability of the mainframe to provide and maintain an acceptable level of service even when things go wrong! Now we know that mainframes don’t crash like they used to in the 1980s, but even so, things can go wrong.

In an ideal world, organizations would take a copy of their data at regular and frequent intervals and restore from the most recent copy in the event of a problem. That would result in only a few minutes of recent changes being lost. It would also create a massive overhead and require a huge amount of storage space. Some companies can afford a hot standby site, which is updated almost as soon as the main site’s data is changed. Should the main site go down, the standby site can take over very quickly and, hopefully, no data is lost.

Other sites take full backups once a week, and incremental backups every evening. That way, it’s possible to restore a file to its state yesterday. If journaling takes place, there will be a file that can be used to restore data almost up to just before the failure.

What I’m illustrating is that a lot of work has gone into getting backups right. What I would also suggest is that not enough attention has been spent on getting the restore part of the operation working as quickly and effectively as possible.

Let’s suppose that one application has somehow had a catastrophic failure. Let’s suppose that the DASD housing the files has died. Where can the recovery files be restored to? Exactly which backup tapes do I need to restore just those files? How quickly can I get hold of them? It’s the orchestration of the recovery operation that needs to take place in software, not in the head of someone who is out of the office that day, or printed on a piece of paper that could be missing from the backup and restore manual.

I wrote recently about Safeguarded Copy on FlashSystem arrays. It creates a security isolated copy of data that can be used in the event of the original data becoming corrupted. In fact, multiple recovery points can be created, which is great. The question is, how can you quickly decide which recovery point backup you want to restore. What software is there available that would speedily work out which recovery point is the one required and make sure that it is restored? Because, in order to speed up the restoration stage, it needs to be done by software orchestration and not trial and error of someone sitting down in front of a screen and seeing which backup is exactly the one they want. I’m not criticizing FlashSystem arrays, I’m just suggesting that the problem with speedy restores is endemic. Everyone worries about backups and they happen all the time. Not enough people are concerned about the restore process because it doesn’t happen (I’m pleased to say!) very often.

To ensure mainframe resiliency, as much effort must be put into simplifying and organizing the restore process as is put into the backup process, so that any mainframe outages last for as little time as possible and no-one – in particular paying customers – notice.

But that’s not all. What happens if nation state or criminal gang bad actors get into your mainframe? Typically, there is a period of time during which hackers raise their security level, exfiltrate useful data, overwrite backups, encrypt data, and then display a ransom demand. Mainframe resiliency also demands that there be some way to identify the early stages of a ransomware attack and stop it spreading further. It also requires that the corrupted files are restored. For this to happen, some kind of software orchestration is required to ensure that the correct (and uncorrupted) backup files are identified, and the data is restored as quickly as possible.

There’s a lot more to mainframe resilience than people might think when they are sitting comfortably after a good lunch discussing the business continuity plan!

Sunday, 15 August 2021

The mainframe and Cloud PC

When I first started working with mainframes, and it was a long time ago, people used to sit in the main office and work on dumb terminals. The mainframe lived in a highly-secure, climate-controlled, part of the building that could only be accessed by people with appropriate key cards. In fact, the majority of people working on the mainframe had no idea what the mainframe looked like. They’d never seen it because they weren’t the chosen few who had been invited into the machine room. For them, it didn’t really exist. They were simply focused on getting their work done. They would come into the office, power up their terminal, and do whatever needed doing. They didn’t know or care about virtual storage or paging or security. They simply did their work. And went home.

How times have changed. Or have they?

The start of August saw the launch of Cloud PC and Windows 365 from Microsoft. The idea is that everything the user wants lives in the cloud – their data, applications, tools, and settings – and they can access it from just about any device they like to use – which could be a laptop, but could also be an Android or Linux device or even an Apple device.

Basically, Azure Virtual Desktop is used to build a virtual machine on top of any other device. And that runs Windows 365 for the user. All the data, applications, etc are stored in the cloud. Users don’t know or care exactly where it is, they simply get on with their work.

It does all seem to be very similar to how mainframers used to work 40 years ago. Everything you need to do your work is stored somewhere, but you don’t know or care where that is. And you simply get on with your work.

Plus ça change, plus c’est la même chose!

It’s not just Microsoft that has recycled this venerable mainframe way of working, Amazon has too. Amazon has its Workspaces Desktop as a Solution (DaaS) product that users might choose. And, of course, Chromebooks have been around for a while. They work on the principle that the operating system is small, the device doesn’t need to have much computing power, and the work takes place in the cloud somewhere.

So, why would you choose Microsoft’s Cloud PC option? Let’s suppose that you are back in the office working, you haven’t completed some major piece of work, so you simply save it and dash to get your train home. On the train, you can get out your tablet (or even your phone) and continue working. And when you get home, you can boot up your home PC and, again, carry on working. You don’t need to borrow a work PC loaded with everything you need to do your job. As long as you have an Internet connection, you can be productive and work on the same desktop environment. Another benefit is, if you leave your laptop on the train, or have it stolen, there is no data on the device. It is all stored in the cloud, so thieves can’t access corporate sensitive data or personal information of clients, etc.

For corporate IT teams, there are also a number of benefits. The first one is budgeting. Rather than buying in new PCs every year or so for staff, they can calculate how much Windows 365 will cost for their staff. If this works out cheaper than buying new devices over a three-year period, they have better control over their budget. There are different sizes of Cloud PC available, and these have different price tags. So, that must be taken into consideration.

Managing Cloud PCs can be done using Endpoint Manager in much the same way that existing physical devices can be managed. And that means corporate security policies can be applied to Cloud PCs as well as real devices. The Endpoint Analytics dashboard allows IT teams to see whether Cloud PC users need more resources allocated to them (or perhaps less). There’s also the Watchdog Service which looks after connectivity. If users become disconnected, alerts are raised, and suggestions made about how to rectify the situation.

I imagine that we’re all fairly familiar with the security on a mainframe, the big question is what kind of security do you get with Windows 365? Firstly, every Cloud PC managed disk is encrypted. Similarly, all data sent over the Internet is encrypted. Data in use isn’t encrypted.

As you might hope in these days of ransomware attacks, multifactor authentication (MFA) is used when someone tries to login. This uses the Azure Active Directory (Azure AD). So, only people passing that test get to login to Windows 365. As mentioned earlier, Endpoint Manager can apply access policies as people try to login.

Lastly, Windows 365 uses a Zero Trust Architecture (ZTA). In the event that perimeter security has failed, it will continually monitor identities, devices, and services that are being used. Should anyone try to access anything unusual or above their security level, ZTA will flag it and alerts will be raised. Again, all data used lives in the cloud.

Certainly, the idea of low power end devices and high power remote devices – whether that’s a mainframe or the cloud – seems like the way things are going for the next little while. To make accessing your mainframe work in that way would probably require it to be able to be accessed from any browser anywhere. I recently discovered that there is a way to do this. If you’re interested, the company is called MainTegrity, its product is called GateWAY z/OS, and you can find out more on its website at http://gatewayzos.com/

The thing about the IT industry is that ideas come and go – and then come back again. Sometimes we have everything on premise, sometimes we have nothing. As always, interesting times!

Sunday, 8 August 2021

The future of the mainframe


With each new version of z/OS, we see two different things brought together and merged into one operating system. They are the pressures placed on the mainframe by the industry – the direction of travel of the mainframe created by market forces – and secondly, the direction of travel that IBM perceive as the best for the future of the mainframe and, obviously being a business, their own future income. These two forces come together every couple of years and manifest themselves in a new version of the operating system. And that’s what happened recently with the announcement of z/OS 2.5, which should be generally available on 30 September.

So, is this version of the operating system like the captain of some huge supertanker trying to change direction in a stormy sea? Or is it more like small sailing boat enjoying the winds and the currents to push in more or less the direction it wants? Maybe a little tacking and steering will be required, but not much? I’ll let you decide.

IBM has said that Version 2.5 of z/OS is “designed to accelerate client adoption of hybrid cloud and AI and drive application modernization projects”. The move to cloud and the growth in the use of various artificial intelligences everywhere seems to be something most companies are completely onboard with. Recognizing that fact, IBM has affirmed that new AI capabilities “are tightly integrated with z/OS workloads, designed to give clients business insights for more informed decision making”. Developers will be able to access a wider choice of AI-based tools, including TensorFlow and IBM Watson Machine Learning for z/OS.

The other thing that organizations are worried about is security. When hackers were disaffected teenagers and nerdy people working alone, problems with security were bad enough, but now criminal gangs and nation-state bad actors are making ransomware into big business along with drugs and people trafficking. It’s an international problem that no-one can ignore. z/OS 2.5 comes with additional security features that expand pervasive encryption to cover additional types of data sets, such as sequential basic format and large format SMS-managed data sets.

Anomaly Mitigation capabilities seem very interesting and can help to prevent ransomware getting on to the system. It uses the predictive failure analysis (PFA) features, which is designed to prevent problems before they occur, as well as, IBM informs us, “Runtime Diagnostics, Workload Manager (WLM), and JES2 to help further detect anomalous behaviour in near real-time, letting clients proactively address potential problems.”

Other new features include:

  • New Java/COBOL interoperability features. IBM says this interoperability “extends existing application programming models with support for parallel 31-bit and 64-bit addressing, simplifying enterprise application modernization”.
  •  z/OS Container Extensions (zCX) will integrate Linux applications and utilities into z/OS,
  • Additional functions for the integration of cloud storage through “transparent cloud tiering” (TCT) and the support of cloud tier support with the “Object Access Method” (OAM) for data transfer in hybrid cloud storage environments. This will simplify central data archiving and backup on the mainframe.

The good news for anyone who has got to install it is that IBM suggests V2.5 is “expected to be faster and easier to install and upgrade, with one client trial demonstrating the ability to install z/OS more than 30 percent faster than compared with IBM z/OS 2.3 and 2.4”. Added to that there’s a new “simplified management experience supplied by streamlined and automated tasks”, which, they suggest, means “specialty skills may not be required”.

For anyone who already has a mainframe, there are plenty of things in the new version that will make life simpler moving forward. And, assuming you have newish hardware, it is very likely that mainframe sites will gradually start upgrading to Version 2.5. The improvements in security may seem slight, but mainframe security is still the best that’s available on any platform. And that may convince some new companies of the value of using such a powerful workhorse for their computing needs. Particularly, if they have already started moving to the cloud, because the mainframe integrates with the cloud so well.

My conclusion is that it does seem this version of z/OS is giving people what they want and not fighting against the general weather conditions out there.

Sunday, 1 August 2021

Battling mainframe ransomware


The idea that mainframes couldn’t be hacked disappeared a long time ago. People are using penetration testing (pentesting) to see how vulnerable their mainframe is to hackers. People are using file integrity monitoring (FIM) software on their mainframes to identify when files are being altered without authorization and to ensure their backups aren’t being modified. And now IBM has announced anti-ransomware Safeguarded Copy for its FlashSystems and on-premises Storage-as-a-Service offerings, with planned public cloud extensions.

So, what do they mean by Safeguarded Copy? It seems the feature automatically creates data copies in point-in-time immutable snapshots that are securely isolated within the system and cannot be accessed or altered by unauthorised users. Organizations can create these protected point-in-time backups of their critical data as frequently as they want, knowing that the process will have a very small impact on resource utilization.

Note: it’s the standard FlashSystem arrays that are being used. There aren’t separate backup target arrays. The idea is to enable the main system to have its own safeguards against ransomware and be able to recover from an attack. Safeguarded Copy allows user to create multiple recovery points for a production volume, which are called Safeguarded Copy backups, and they are stored in a storage space that is called Safeguarded Copy backup capacity.

Although the data copies created by Safeguarded Copy are security isolated within the systems and cannot be accessed, they are available should normal operations be disrupted by a data breach or cyberattack. And then the copies can be used to recover quickly. You might be wondering how this can be done if the backups aren’t directly accessible by a host. The answer is that the data can accessed once it has been recovered to a separate recovery volume.

The practicalities are that storage administrators can schedule automatic snapshots, which are then stored into safeguarded pools on the storage system. The data has to be recovered (as mentioned above) to become usable. In addition to validating copies of data, the Safeguarded Copy can be used to diagnose production issues.

By integrating Safeguarded Copy with IBM Security QRadar platform for security monitoring, it’s possible for QRadar to look out for signs of a ransomware attacks and proactively trigger Safeguarded Copy to create backups, which can then be used to restore data in the event of a successful attack.

With the IBM and Ponemon Cost of a Data Breach Report 2020 showing that the average total cost of a data breach was $3.86 million, and that figure went up to $8.64 million for organizations based in the USA, it makes sense for IBM to make security a top priority in their development work. There is much discussion at the moment about whether companies should be obliged to reveal not only whether they’ve paid a ransom, but how much they paid. I’m sure that when those figures are fully revealed, it will start to make sense to the accountants at most organizations to spend money wisely beforehand to ensure that they are not funding hackers – who could be criminal gangs or nation state actors – after a ransom has been received, their data has been sold on the dark web, and their reputation has been muddied.

A report in March from Palo Alto Networks found that the average payment following a ransomware attack in 2020 was up 171 percent to $312,493, compared to $115,123 in 2019. The report also found that the highest ransom demanded in 2020 was $30 million, which was double the highest of $15 million during 2015-2019. The largest payout that the survey found was $10 million.

This just adds weight to the argument for mainframe-using organizations to spend some money up front, whether that is on pentesting, FIM software, or Safeguarded Copy on FlashSystems, or anything else that works, to prevent successful ransomware attacks happening to them.