Sunday, 23 February 2025

Mainframes, AI, and security

I was interested to read IBM’s thoughts about AI on the mainframe, which was published in January. You can read it here. The article discusses the different ways that AI can be integrated with mainframes. The article tells us that on-chip AI accelerators can scale and process millions of inference requests per second at very low latency rates. This capability allows organizations to use data and transactional gravity by strategically co-locating large datasets, AI, and critical business applications. In the future, next-gen accelerators will open up new opportunities to expand AI capabilities and use cases as an organization’s needs grow.

It talks about ensemble AI, which it describes as a hybrid concept that integrates different AI technologies, such as traditional AI and LLM encoder models, to deliver faster, more accurate results than any single model can accomplish alone, tapping into the mainframe's massive processing power and data storage capabilities.

The article then discusses four potential use cases of AI on a mainframe. The first of these is real-time fraud detection, which can be of use to fintech companies. As an example, it discusses a large North American bank that had developed an AI-powered credit-scoring model and deployed it on an on-premises cloud platform to help fight fraud. However, only 20% of credit card transactions could be scored in real-time. The bank decided to move the complex fraud-detecting tools to its mainframe.

After the mainframe implementation, the bank began scoring 100% of credit card transactions in real-time, with 15,000 transactions per second, providing significant fraud detection.

Moreover, each transaction used to take 80 milliseconds to score. With the reduced latency provided by the mainframe, response times now occur in 2 milliseconds or less. This move to the mainframe has also saved the bank over US$20 million in annual fraud prevention spend without impacting service-level agreements.

The second example was IT operations and AIOps describing how organizations can now use AI to proactively prevent or even predict an outage caused by equipment failure. By applying AI mechanisms, organizations can detect anomalies at the transaction, application, subsystem, and system levels. For instance, sensors can analyse data from mainframe components to predict potential hardware failures and enable preventative maintenance. They say that organizations are increasingly turning to the application of AI capabilities to automate, streamline, and optimize IT infrastructure and operational workflows. AIOps enables IT operations teams to respond quickly to slowdowns and outages, providing better visibility and context.

The third example given is advanced document processing, saying that processing documents on the mainframe helps streamline and deliver accurate data extraction in a highly secure setting. Organizations can use gen AI to summarize financial documents and business reports, extract key data points (for example, financial metrics and performance indicators), and identify essential information for compliance processes (for example, financial audits).

And lastly in their list are AI code assistants. They affirm that virtual assistants on the mainframe are helping to bridge the developer skill gap. Tools, such as IBM® watsonx Code Assistant™ for Z, use generative AI to analyse, understand and modernize existing COBOL applications. This capability allows developers to translate COBOL code into languages like Java. It also accelerates application modernization while preserving legacy COBOL systems' functionality. Watsonx Code Assistant for Z features include code explanation, automated refactoring and code optimization advice, making it easier for developers to maintain and update old COBOL applications.

Now I’m not saying anything against those four areas. In fact, I totally support them as great uses of AI on a mainframe. However, I would have thought that one area where AI assistance would be needed is in security. It only takes a brief Google search to find a number of companies that have produced reports about ransomware attacks or giving more details about the techniques criminal gangs or teams associated with foreign governments are using to attack organizations. There are also plenty of reports about the cost of these attacks to more high-profile organizations. I don’t just mean the cost of new hardware, software, or staff, I mean fines for non-compliance with regulations, and court costs and fines paid to individuals whose data has been stolen.

I would have thought those kinds of stories would have crossed the desk of an organization’s chief financial officer (CFO) as well as anyone associated with IT. Admittedly, the majority of attacks are on non-mainframe platforms, but that doesn’t mean mainframes aren’t targets for attacks because, as we know, they contain a large amount of data about people and finances.

I would like to see AI-based software able to be as effective as the best non-AI security software. And then I would like to see the AI software learn and improve. As I’ve mentioned previously, the security software needs to be trained to recognize ‘normal’ activity by people who have access to the mainframe, and then automatically suspend any unusual actions by them. This prevents too much damage being done, if it a job being run by malware rather than a real person. If the person is authorized, then appropriate checks by the security team can allow the job to continue.

Because malware attacks get more sophisticated each year, it’s important to have some kind of defence shield that can learn an adapt and continue to keep the mainframe safe. I’m surprised that we haven’t had security software listed as an important area for AI development. I assume it must be because it’s not such an easy thing to do as some of the other areas listed in the IBM article.

Sunday, 16 February 2025

Trevor Eddolls – IBM Champion 2025

The iTech-Ed Group is pleased to announce that Trevor Eddolls, its Chair, has been recognized by IBM as an IBM Champion for 2025. Trevor was first awarded IBM Champion status in 2009.

A blue square with white text

AI-generated content may be incorrect.


 

IBM said: “On behalf of IBM, it is my great pleasure to recognize you as a returning IBM Champion in 2025. Congratulations!

“We would like to thank you for your continued leadership and contributions to the IBM technology community. This recognition is awarded based on your renewal and contributions for the 2024 calendar year. The IBM Champion designation is a 1-year term, and may be renewed by IBM annually, provided you demonstrate continued community engagement and contributions. You may also have earned 2024 IBM Rising Champion advocacy badges (IBM Contributor, Advocate, and Influencer) on your way to this honour. Your IBM Champion status renews now and will run through December 2025.”

Trevor Eddolls, Chair of iTech-Ed Ltd said: “I think it's really important in these days of multiple computing platforms being available that people share information with others about the positive contributions mainframes make to the world of IT. And I'm proud that my efforts have been recognized again this year by IBM. I think the Champion programme is a very positive way for IBM to recognize people around the world who help to promote its products and share their skills in using them.”

According to IBM: “The IBM Champion program recognizes these innovative thought leaders in the technical community and rewards these contributions by amplifying their voice and increasing their sphere of influence. IBM Champions are enthusiasts and advocates: IT professionals, business leaders, developers, executives, educators, and influencers who support and mentor others to help them get the most out of IBM software, solutions, and services.”

So why is iTech-Ed Group’s Trevor Eddolls an IBM Champion? Well, he doesn’t work for IBM, but he does write about mainframe hardware and software. You can read his articles here. He has also written articles for the TechChannel website, and often blogs on the Planet Mainframe website. Trevor has spoken at the GSE UK regional conference for the past few years. In 2024, he was talking about how to create artificial generalized intelligence (AGI). He has been Editorial Director for the well-respected Arcati Mainframe Yearbook (renamed the Arcati Mainframe Navigator in 2025). And Trevor Eddolls was the chair of the Virtual IMS, the Virtual CICS, and the Virtual Db2 user groups until recently. Their new website can be found at virtualusergroups.com. And this work has earned Trevor Eddolls the IBM Champion accolade for the past seventeen years.

Are IBM Champions compensated for their role? No. Do IBM Champions have any obligations to IBM? Again, the answer is no. The title recognizes their past contributions to the community only over the previous 12 months. Do IBM Champions have any formal relationship with IBM? No. IBM Champions don’t formally represent IBM, nor do they speak on behalf of IBM.

But it’s not all one-sided! There are regular IBM Champions calls, where IBM and Champions share relevant information on a range of topics. IBM Champions also receive merchandise customized with the IBM Champion logo. And IBM Champions receive visibility, recognition, and networking opportunities at IBM events and conferences; and special access to product development teams, and invitations and discounts to events and conference.

You can find more information about the Trevor and his work on X (Twitter), FacebookInstagram, and LinkedIn.

You can read Trevor's IBM Champion profile here.

You can find out more about iTech-Ed here.

 

Sunday, 2 February 2025

Mainframe staff, security, and AI

If you want to test a new application, the best data to test it on is live data! Now, I’m sure that there are procedures in place to not do that. I’m sure anonymized data would be used instead. But it became apparent a few years ago that some members of staff were copying live data off the mainframe and testing it in cloud applications. Again, hopefully this doesn’t happen anymore. However, there is apparently a new problem facing mainframe security teams. And that is using live data on artificial intelligence (AI) applications.

It was the rapid increase in people working from home during the pandemic that led to a rise in shadow IT – people using applications to get work done, but those applications hadn’t been tested by the IT security team. A recent survey has found that AI is now giving rise to another massive security issue. This becomes even more of an issue with the current popularity of Deepseek V3, and the announcement of Alibaba Qwen 2.5, both AIs originating from China.

Cybsafe’s The Annual Cybersecurity Attitudes and Behaviors Report 2024-2025 found that, worryingly, almost 2 in 5 (38%) professionals have admitted to sharing personal data with AI platforms, without their employer’s permission. So, what data is being shared most often and what are the implications? That’s what application security SaaS company, Indusface, looked into. Here’s what they found.

One of the most common categories of information shared with AI is work-related files and documents. Over 80% of professionals in Fortune 500 enterprises use AI tools, such as ChatGPT, to assist with tasks such as analysing numbers, refining emails, reports, and presentations.2

However, 11% of the data employees paste into ChatGPT is strictly confidential, for example internal business strategies, and the employees don’t fully understand how the platform processes this data. Staff should remove sensitive data when inputting search commands into AI tools.3

Personal details such as names, addresses, and contact information are often being shared with AI tools daily. Shockingly, 30% of professionals believe that protecting their personal data isn’t worth the effort, which indicates a growing sense of helplessness and lack of training.

Access to cybersecurity training has increased for the first time in four years, with 1 in 3 (33%) participants using it and 11% having access but not utilizing it. For businesses to remain safe from cyber security threats, it is important to carry out cybersecurity training for staff, upskilling on the safe use of AI.1

Client information, including data that may fall under regulatory or confidentiality requirements, is often being shared with AI by professionals.

For business owners or managers using AI for employee information, it is important to be wary of sharing bank account details, payroll, addresses, or even performance reviews because this can violate contract policy and lead to organization vulnerability due to any potential legal actions if sensitive employee data is leaked.

Large language models (LLMs) are often used and are crucial AI models for many generative AI applications, such as virtual assistants and conversational AI chatbots. This can often be used with Open AI models, Google Cloud AI, and many more.

However, the data that helps train LLMs is usually sourced by web crawlers scraping and collecting information from websites. This data is often obtained without users’ consent and might contain personally identifiable information (PII).

Other AI systems that deliver tailored customer experiences might collect personal data, too. It is recommended to ensure that the devices used when interacting with LLMs are secure, with full antivirus protection to safeguard information before it is shared, especially when dealing with sensitive business financial information.

AI models are designed to provide insights, but not safely secure passwords, and could result in unintended exposure, especially if the platform does not have strict privacy and security measures.

Indusface recommends that individuals avoid reusing passwords that may have been used across multiple sites because this could lead to a breach on multiple accounts. The importance of using strong passwords with multiple symbols and numbers has never been more important, in addition to activating two-factor identification to secure accounts and mitigate the risk of cyberattacks.

Developers and employees increasingly turn to AI for coding assistance, however sharing company codebases can pose a major security risk because it is a business’s core intellectual property. If proprietary source code is pasted into AI platforms, it may be stored, processed, or even used to train future AI models, potentially exposing trade secrets to external entities.

Businesses should, therefore, implement strict AI usage policies to ensure sensitive code remains protected and never shared externally. Additionally, using self-hosted AI models or secure, company-approved AI tools can help mitigate the risks of leaking intellectual property.

 

 

The sources given for their research are:

  1. Cybsafe | The Annual CybersecurityAttitudes and Behaviors Report 2024-2025
  2. Masterofcode | MOCG Picks: 10 ChatGPT Statistics Every Business Leader Should Know
  3. CyberHaven | 11% of data employees paste into ChatGPT is confidential

 

Sunday, 19 January 2025

AI and ethics and mainframes

Imagine two people talking in a bar and one says that they believe in God and the other says that there is no such thing. The conversation moves on. One says that they think their Apple phone and tablet are the best things ever and the other says that if most of the world uses Android that must prove their thinking is wrong. The conversation moves on. One person says that Trump is the best person to lead the USA into the future and the other says that Trump will only harm the country’s standing in the world.

It doesn’t matter which person you identify with in each of those discussions, what it shows is that not all people agree on these three and many other issues. But we knew that already. The reason why it is important is because those two hypothetical people could be responsible for the training of two different pieces of artificial intelligence (AI) software. The views, opinions, beliefs, and values of the person responsible for the training of an AI could influence the ‘thinking’ of the AI and the responses that it comes up with when asked questions by users. And those people could be mainframe users.

Britannica tells us that the “term ethics may refer to the philosophical study of the concepts of moral right and wrong and moral good and bad, to any philosophical theory of what is morally right and wrong or morally good and bad, and to any system or code of moral rules, principles, or values”.

Let’s suppose that someone with the mindset and ethics of Adolf Hitler trained a popular AI, or perhaps one of the founding fathers of the USA was responsible for the training. What kind of AI would they produce. The founding fathers of the USA were generally quite happy with the idea of slavery. The men still thought that women didn’t need to be educated because their poor feeble female brains couldn’t cope. And that women basically belonged to their fathers until they were married when ownership passed to their husbands. Much the same thinking applied over most of Europe. It’s the way most Europeans thought in the 17th and 18th centuries.

So, let’s suppose that a piece of AI software – and, nowadays, you can hardly buy a new device without it being advertised as coming with some super new AI – has been trained with some ethical value that the majority of people don’t agree with. However, because that is such a small part and everything else seems OK, the software gets installed on your mainframe. Let’s suppose it’s a piece of security software that is identifying unusual activity on your mainframe. Perhaps a systems programmer has apparently logged in from a foreign country at 2am, and is now making changes to the system. Perhaps he is giving some software higher access levels than before. Perhaps he is deleting certain files. Now, hopefully, your security AI will spot this as unusual, and quickly suspend the job until someone can check exactly what is happening. Then, if it’s all OK, the job can continue. If it’s not OK, then not too much damage has been done.

The users of the AI will assume that the AI is on their side. It has the same values as them and knows what’s good and bad, or right and wrong in the same way as the user. But what if it doesn’t? You don’t usually expect software to have ethical values, but with AI, this becomes more of a concern. What about using an open-source AI. How can you check whether the values that have been trained into it match yours?

There’s lots of talk about the ethics of using AI software. Should students use AI to write their essays. Should AI be used to create nude videos of famous (and not so famous) people. And there are so many other areas. But what no-one talks about is the actual ethical values of the AI software itself.

We’re all familiar with the Terminator movies. Suppose the AI decides that humans are destroying the planet, and the right thing is to remove them from existence. Or, more worryingly, suppose the AI decides that someone logging into your mainframe from one specific foreign country is permissible because they are our friends, and lets them launch a ransomware attack on your mainframe.

Ethical conversations over a beer usually pass off without anyone getting too upset. The embedded ethics of AI software might have more far-reaching consequences.


Sunday, 12 January 2025

2024 at iTech-Ed Ltd

As usual at this time of year, I thought I’d take a look at the previous year, with the spotlight on what was happening at iTech-Ed Ltd.

The exciting news in January was that Trevor Eddolls was recognized by IBM as a 2024 IBM Champion. IBM said: “On behalf of IBM, it is my great pleasure to recognize you as a returning IBM Champion in 2024. Congratulations! We would like to thank you for your continued leadership and contributions to the IBM technology community. This recognition is awarded based on your contributions for the 2023 calendar year.”

On 16 January, the Virtual Db2 user group saw a presentation from Marcus Davage, Lead Product Developer at BMC Software. He was discussing how “Driving Down Database Development Dollars”. Then on 23 January, Todd Havekost, Senior z/OS Performance Consultant at IntelliMagic discussed “Enhanced Analysis Through Integrating CICS and Other Types of SMF Data” with the Virtual CICS user group.

February saw the publication of the always popular Arcati Mainframe Yearbook 2024. You can download a copy here – it’s FREE. Last year’s edition of this highly-respected annual source of mainframe information was downloaded around 21,000 times during the course of the year.

Also in February, Trevor’s article, “The Comprehensive Beginners’ Guide to AI”, was published on the TechChannel website. And his article “Ransomware Attacks and your Health” was published on the Planet Mainframe website.

On 13 February, the Virtual IMS user group had a presentation from Dr Daniela Schilling, CEO of Delta Software Technology, entitled “Replacing IBM IMS DB – Fully Automated and with Highest Security”.

On 12 March, Jenny He PhD, IBM Master Inventor, CICS Development at IBM Hursley Park, gave a presentation to the Virtual CICS user group entitled, “CICS Event processing and CICS policies”. And on 19 March, Toine Michielse, Solutions Architect at Broadcom, discussed “A day in the life of a Db2 for z/OS Schema” with the Virtual DB2 user group.

In April, Trevor Eddolls was awarded an IBM Z and LinuxONE Community Contributor – 2024 (Level 1) badge. The badge earner is an external community member who is passionate about IBM zSystems and LinuxONE and wants to make a positive difference. This individual has expressed interest in contributing to the community in their own unique way.


Also in April, Trevor’s article, “Why Today’s AI Is Failing”, was published on the TechChannel website.

On 9 April, the Virtual IMS user group had a presentation from IBM’s Stephen P Nathan entitled “How to Help IBM AND YOU Quickly Resolve IMS Problems”.

Towards the end of April, Trevor was listed as a 2024 Influential Mainframers on the Planet Mainframe website.

In July, Trevor Eddolls was awarded an IBM Z and LinuxONE Community Advocate – 2024 (Level 2) badge. The badge earner has actively contributed to the IBM Z and LinuxONE community and has expressed interest in continuing to do so. This individual is in good standing with their IBM Z and LinuxONE peers and is passionate about taking their advocacy to the next level. The badge earner has expert skills in IBM Z and LinuxONE and can be expected to regularly contribute technical knowledge to the community.

Also in July, Trevor Eddolls was awarded an IBM Z and LinuxONE Community Influencer – 2024 (Level 3) badge. The badge earner is an active and passionate member of the IBM Z and LinuxONE Community. He is a thought leader and viewed as a technical expert by his peers. This individual contributes to the community regularly.

iTech-Ed Ltd was shortlisted in the seventh annual Southern Enterprise Awards. The team at SME News nominated iTech-Ed Ltd, recognizing its exceptional contributions and achievements.

In August, Trevor’s article, “Get ready for DORA”, was published on the Planet Mainframe website.

Also in August, iTech-Ed Ltd was awarded, "Best Specialist IT Consultancy 2024 - Wiltshire" in the Southern Enterprise Awards 2024, hosted by SME news. Also, C Level Focus emailed to say, "You've been named one of the 'Top 10 Inspiring CEOs of 2024' by the CLF Magazine editorial team".

Thirdly in August, Mainframerz Meetup on LInkedIn said...

Perhaps not the best way to deal with a data leak

For many years
Trevor Eddolls has written many great posts and I couldn't think of a better way to lead into the IBM Cost of a Data Breach report that follows than the story by Trevor of a data breach at NTT.
If you would like to read a 'how to guide' of how not to respond to a Data Breach then the article from Trevor is 'the guide you are looking for' – (you need to read the quote in a Star Wars voice).
I genuinely think you may take an in-take of breath in how this was responded to, I don't want to give the best bits away and it's all very juicy. I will leave one teaser which is one of the statements shared, which the politest way of saying this would be that this is not full disclosure of the facts.
Additional measures to mitigate any further risk and protect the data of our customers were also activated. At this time, there is no visibility that client data has been affected.
If there was to be a scoop of the year I vote for this article Trevor has shared as it is eye opening and a must read recommendation. You can read the full shenanigans that Trevor has reported on
here.

In October, Trevor’s article, “Ransomware isn’t really a problem, is it?”, was published on the Planet Mainframe website.

Also in October, Trevor Eddolls was awarded the IBM Contributor, Advocate, and Influencer – 2024 badges.


At the end of October, Trevor’s article, “Defense against the dark arts — mainframe security”, was published on the Planet Mainframe website.

On 5 November, Trevor presented to the AI stream at the GSE conference in the UK. His presentation was called “How to create Artificial Generalized Intelligence”, and looked at how the human brain has solved the problem of a generalized intelligence and how this can be applied to AI.

In December Trevor Eddolls was awarded a speaker badge. The award says, at "Mainframe@60: The Diamond Anniversary of Digital Dominance", you showcased profound expertise and in-depth knowledge. Your engaging presentation style and ability to foster interactive discussions left a lasting impression on the participants, making your session a valuable and enriching experience to our conference.”

Looking forward to the coming year, the Arcati Mainframe Yearbook 2025 will be published under its new title of the Arcati Mainframe Navigator. The Virtual IMS, Virtual CICS, and Virtual Db2 user groups will continue to meet six times a year. All of those are now curated by the great team at Planet Mainframe. And who knows what else we have to look forward to.

 

Sunday, 8 December 2024

Cyber targets for 2025

Let us imagine that there is a room somewhere in Russia (but it could be anywhere else hostile to the West) and it’s full of hackers plotting their attacks for 2025. You can imagine that they are sharing stories of their successes in 2024. How they have targeted people with phishing emails and got them to open malware or download (unwittingly) malware that has not only given the hackers access to the servers of that company, but every other company in the supply chain.

The next hacker speaks up explaining how he has got round the security of cloud providers and managed to get into a variety of organizations that way. He proudly explains that he hasn’t even exploited some of those hacks yet. They are now easy targets for the New Year.

A third hacker explains how he managed to access a security update to a frequently used piece of software, and how he had added a back door that no-one had spotted. So, when everyone downloaded the software and patched the vulnerability, they introduced a back door that only he knew about. He suggested that this time next year he would be rich from all the ransoms he was going to collect.

Another hacker jumps up and explains that he was using AI to automate ransomware attacks, and he is making lots of dosh from the people who were paying him for the Ransomware as a Service software – sometimes people with very little IT knowledge – and were then using it to attack companies that had upset them in some way.

Lots of other people want to speak up with stories of how they had attacked companies and made money, but everyone stops speaking as an old general gets to his feet. He looks very stern but smiles as he starts to speak. “Comrades”, he says, “you have all done very well attacking companies in the West.” He pauses and his face takes on a sternness that had scared many a junior officer. He continues, “The problem is this: we have not defeated the West. What I need you to do is find some way to bring down the whole infrastructure of western society. Can you do that?”

The hackers look round at each other, until one speaks up. “Capitalist society depends on capital.” The audience is not overimpressed by the obviousness of the comment. There is much murmuring from the audience, but the hacker continues, “Why don’t we attack the banks and all the other financial institutions in North America and Europe. If they don’t have access to money, everything else will come to a stop.” The crowd nods in agreement. Some make additional useful comments to each other.

“How do we do that?” asks the general. “We attack the mainframes that are used by most of these organizations”, replies the hacker. And that’s what they do. Attacks by people who understood Windows and Linux continue in all their forms, but a large tranche of the technical people are given the job of understanding how mainframes work and their vulnerabilities. After all, the majority of financial institutions use mainframes. A subgroup is given the task of looking at employees on mainframes and seeing which ones could be manipulated into giving access to these fintech mainframes. They are looking for staff with drug habits and staff with financial problems or other issues that could be used against them. Another group has the task of getting keyloggers onto the laptops of systems programmers at mainframe sites.

A list of potential hacking techniques that have been used before are circulated amongst the hackers for them to see which still work and are useful for others to try.

They could attack sites using CICS. There are automated tools like CICSpwn available that could be used to identify potential misconfigurations, which could then be used by the hackers to bypass authentication. They could use the CICS customer front end and try a simple brute force attack to find a userid and password that would get them into the system.

They could use FTP. Two things need to happen first – keylogger software needs to capture the login credentials from a systems programmer, and a ‘connection getter’ needs to identify where to FTP to. Commands can be written to upload malicious binaries, and JES/FTP commands can be used to execute those binaries.

They could use TN3270 emulation software for their attack. Provided they have some potential userids, they could try password spraying, ie a few commonly-used passwords can be tried against every userid on the system.

NJE allows one trusted mainframe to send a job to another mainframe that it’s connected to. Hackers could use NJE to spoof a mainframe or submit a job and gain access to that other mainframe.

Then there’s potential vulnerabilities in Linux and other non-IBM software (like Ansible, Java, etc) that runs on mainframes.

Other techniques are available, but it’s not the function of this blog to make the job of nation state hackers easier. It is the job of this blog to ensure that every mainframe site is doing everything it can to ensure that it is secure against all forms of attack, and that it has software installed that can alert staff at the earliest opportunity that an attack has started, and the defence software needs to be able to suspend any suspect jobs as soon as possible.

Meanwhile, meetings like the one I’ve envisaged are probably going on, and mainframe-using companies in the West are going to be the targets in 2025. Don’t let yours be one of them.

Sunday, 1 December 2024

Rock solid AI – Granite on a mainframe

Let’s start with what people are familiar with, ChatGPT. ChatGPT is a highly-trained and clever chatbot. The GPT part of its name stands for Generative Pre-trained Transformer. Generative means that it can generate text or other forms of output. Pre-trained means that it has been trained on a large dataset. And Transformer refers to a type of neural network architecture enabling it to understand the relationships and dependencies between words in a piece of text. IBM’s Granite 3.0 is very similar to ChatGPT, except that it is optimized for specific enterprise applications rather than general queries.

Just a side note, I was wondering about the choice of name for the product. In the UK, the traditional gift for a 90th anniversary is granite. I just wondered whether there was some kind of link. In 1933 IBM bought Electromatic Typewriters, but I can’t see the link. Or maybe I’ve been doing too many brain-training quizzes!

Granite was originally developed by IBM and intended for use on Watsonx along with other models. In May this year, IBM released the source code of four variations of Granite Code Models under Apache 2, allowing completely free use, modification, and sharing of the software.

In the original press release in September 2023, IBM said: “Recognizing that a single model will not fit the unique needs of every business use case, the Granite models are being developed in different sizes. These IBM models – built on a decoder-only architecture – aim to help businesses scale AI. For instance, businesses can use them to apply retrieval augmented generation for searching enterprise knowledge bases to generate tailored responses to customer inquiries; use summarization to condense long-form content – like contracts or call transcripts – into short descriptions; and deploy insight extraction and classification to determine factors like customer sentiment.”

The two sizes mentioned in that press release are the 8B and 2B models.

In October this year, Version 3.0 was released, which is made up of a number of models. In fact the press release tells us that “IBM Granite 3.0 release comprises: 

  • Dense, general purpose LLMs: Granite-3.0-8B-Instruct, Granite-3.0-8B-Base, Granite-3.0-2B-Instruct and Granite-3.0-2B-Base.
  • LLM-based input-output guardrail models: Granite-Guardian-3.0-8B, Granite-Guardian-3.0-2B
  • Mixture of experts (MoE) models for minimum latency: Granite-3.0-3B-A800M-Instruct, Granite-3.0-1B-A400M-Instruct
  • Speculative decoder for increased inference speed and efficiency: Granite-3.0-8B-Instruct-Accelerator.

Let’s put a little more flesh on the bones of those models:

  • The base and instruction-tuned language models are designed for agentic workflows, Retrieval Augmented Generation (RAG), text summarization, text analytics and extraction, classification, and content generation.
  • The decoder-only models are designed for code generative tasks, including code generation, code explanation, and code editing, and are trained with code written in 116 programming languages.
  • The time series models are lightweight and pre-trained for time-series forecasting, and are optimized to run efficiently across a range of hardware configurations.
  • Granite Guardian can safeguard AI by ensuring enterprise data security and mitigating risks across a variety of user prompts and LLM responses.
  • Granite for geospatial data is an AI Foundation Model for Earth Observations created by NASA and IBM. It uses large-scale satellite and remote sensing data.

In case you didn’t know, agentic workflows refer to autonomous AI agents dynamically interacting with large language models (LLMs) to complete complex tasks and produce outputs that are orchestrated as part of a larger end-to-end business process automation.

Users can deploy open-source Granite models in production with Red Hat Enterprise Linux AI and watsonx, at scale. Users can build faster with capabilities such as tool-calling, 12 languages, multi-modal adaptors (coming soon), and more, IBM tells us.

IBM is claiming that Granite 3.0 is cheaper to use compared to previous versions and other LLM (large language models) such as GPT-4 and Llama

IBM also tested the Granite Guardian against other guardrail models in terms of their ability to detect and avoid harmful information, violence, explicit content, substance abuse, and personal identifying information, showing it made AI applications safer and more trusted.

We’re told that the Granite code models range from 3 billion to 34 billion parameters and have been trained on 116 programming languages and 3 to 4 terabytes of tokens, combining extensive code data and natural language datasets. If you want to get your hands on them, the models are available from Hugging Face, GitHub, Watsonx.ai, and Red Hat Enterprise Linux (RHEL) AI. A curated set of the Granite 3.0 models can be found on Ollama and Replicate.

At the same time, IBM released a new version of watsonx Code Assistant for application development. The product leverages Granite models to augment developer skill sets, simplifying and automating their development and modernization efforts. It simplifies and accelerates coding workflows across Python, Java, C, C++, Go, JavaScript, Typescript and more.

Users can download the IBM Granite.Code (which is part of the watsonx Code Assistant product portfolio) extension for Visual Studio Code to unlock the full potential of the Granite code model from here.

It seems to me that the Granite product line is a great way for organizations to make use of AI both on and off the mainframe. I’m looking forward to seeing what they announce with Granite 4.0 and other future versions.