Co-founder of Strumenta. Essays on software languages, legacy systems and the craft of building tools.

275 articles · since 2009

Rows of wooden barrels in a brewery storehouse
Legacy systems & modernization · 9 min · 1,887 words · 1 of 34 in this topic · by

Codebase Blueprint: it all started with Legacy Codebases

As part of our job we have to understand quickly legacy codebases. We dive into codebases that have decades of history and possibly millions of lines of code and we need to understand as much as possible about them as quickly as we can. To do that we have an ace up our sleeve, and this is what I am going to show in this article.

This happens to us because we help companies migrate their legacy systems and this starts with doing a plan, but how can one do a plan if one does not understand the codebase first? We do the migration planning bit through our Migration Blueprint service. A Migration Blueprint typically takes between 10 and 13 weeks. Understanding a codebase of that size, in that time, well enough to give advice somebody will act on, is a challenge. By that I mean, a very difficult challenge.

So we built a tool to help with that. The tool morphed over the years and went through different names. We used to call it Code Insight Studio, and you can find it mentioned under that name in the articles where we described the Migration Blueprint and the Migration CodeCraft. Now we call it Codebase Blueprint. Given we are engineers and creativity is not our strong point, the name of the tool comes from the name of the service. I like to think that we are much better at Language Engineering than at naming. If that is not the case please do not tell me.

The tool is built on the same technology we use for all our other systems: Starlasu and the parsers we developed over the years. This means it understands many different languages: RPG, COBOL, DDS, SQL in its different dialects, PL/SQL, SAS, Visual Basic 6, and many others. On top of that, it has visualizations that are specific to certain environments. For example, for RPG we can render the green screens defined in display files. For Visual Basic 6 we can render the forms.

In this article I want to show you this tool. At the end I will share a few thoughts on where I think tools like this one will be needed in the future.

The video

We recorded a demo of the tool. You can watch it here, or keep reading: in the rest of the article I go through the same features.

The example: TOBi

We cannot show the code of our clients, unless we want them to sue us, and half my job is to avoid that. So we used TOBi, a small open-source IBM i application published by IBM. It is an order management system with the usual stuff: customers, articles, orders, and VAT.

It is small, about 100 files while real ones have… a bit more. That said, I think it has a realistic mix of bits and parts as it contains RPG III, RPG IV both in fixed format and free format, embedded SQL, display files, physical and logical files, and one printer file.

Overview

When we open a codebase, we land on the overview.

The overview of TOBi in Codebase Blueprint: 103 files, 8,994 lines, 1,515 resolved cross-file references

For TOBi we get 103 files and 8,994 lines: 45 programs, 12 copybooks, 18 display files and 26 other DDS files. They are connected by 1,515 cross-file references. The strongest one goes from the screen PRO201D to the provider file: 48 references.

These references are identified through static analysis. This means: deterministic and reliable.

Statistics

The Statistics tab: 21,100 syntax nodes and 225 kinds of construct, with the most used statements and expressions

TOBi has 21,100 nodes in its syntax trees, of 225 different kinds. For each kind we see how many times it is used and in how many files.

This is what we look at first in a migration. Every client uses their own subset of the language. A statement used in 35 files must be supported, and supported well. A statement that appears in a single file, because one programmer wanted to try it one day he was bored, well, perhaps we can give it a little less attention.

The tool also lists migration indicators: constructs that typically need attention. We may want to know for example about how people have been using GOTOs.

The map

The map shows how the files depend on each other. By default it groups them by folder.

The architecture map grouped by folder: QRPGLESRC points to QDDSSRC with 763 references

The programs in QRPGLESRC refer to the data definitions in QDDSSRC 763 times. Not a surprise. The DDS files then point to common, which contains a single file: SAMREF, the field reference file, where field sizes are defined once.

Folders are a good thing but one could also group files logically:

The same map grouped into 14 detected modules

The tool groups files that are tightly related into modules: 14 in this case. Each module mixes programs, screens and database files. Modules give us the order of migration: we can migrate and test one module at a time, instead of doing one big, risky release.

We identify modules by finding clusters of relationships. Pretty neat, eh?

Following a change request

Now let’s use the tool for a real task. The business wants longer customer names: from 30 to 50 characters. What do we have to touch?

1. Where the size is defined

SAMREF.PF, line 21: CUSTNM 30, TEXT('CUSTOMER NAME')

The size is in SAMREF, line 21: CUSTNM 30. The customer file does not repeat it. It just says CUSTNM R, a reference field, and the tool resolves that reference to SAMREF.

2. Who uses the customer file

The Relationships view centred on CUSTOMER.PF: 12 inbound files and 1 outbound (SAMREF)

The Relationships view shows the neighbourhood of CUSTOMER.PF. Twelve files point to it, and it points to one file: SAMREF. Each arrow carries the number of references behind it, and we can click it to see them.

3. Who uses the field

We do not care about the whole file, just about one field. So in the References panel we filter by CUSTNM:

References to CUSTOMER.PF filtered on CUSTNM: 12 resolved references

12 references in 11 files: 5 display files, the logical file CUSTOME2, the printer file ORD500O and 4 programs. CUSTOME2 is keyed on the customer name, so its access path changes too.

This is our list of places to check.

4. The screen

We open the first one, CUS200D, and switch to the Screen view.

The CUS200D display file rendered as a 5250 screen, with the Customer column selected

The screen is drawn from the DDS source. No emulator, no running system, no LLM guessing. Clicking the Customer column selects line 17 in the source.

And here we have a problem. The name starts at column 13, the city at column 44. At 50 characters the name runs over the city. This screen needs a new layout, not just a recompile.

5. Where the program reads the name

CUS300 is the customer service module. Its procedure GetCusName chains the customer by ID and returns CUSTNM. We put the cursor on CUSTNM and press Shift+F12:

Usages of CUSTNM in CUS300: defined at line 3, used at lines 17 and 22

The field is not declared in the program. It comes from the file declared on line 3. Pressing F12 from there we reach the logical file CUSTOME1, and from it the physical file CUSTOMER. From the program to the database, one step at a time.

Notice also the small counters in the code, like [14 uses] on chainCUSTOME1.

6. A hard-coded length

The order program ORD100 gets the name by calling GetCusName. F12 on the call takes us to the prototype, in a copybook:

The prototype of GetCusName in CUSTOMER.RPGLEINC: PR 30

PR 30. The prototype returns 30 characters, hard-coded. It does not come from SAMREF. If we change only SAMREF, the name gets silently truncated here.

No reference field would have told us that. We found it by following the code.

7. What happens when an order is confirmed

A view I personally like is the sequence diagram. This is subroutine s01act in ORD100:

The sequence diagram of subroutine s01act in ORD100, generated from the code

It is generated from the code. If the order is confirmed, we write the order, read the temporary lines and write each detail, then call prtOrd. If the option is 4, we delete the temporary lines and update the subfile. Clicking a message opens the corresponding statement.

Years ago we wrote about transforming RPG code into sequence diagrams, using our RPG parser and PlantUML. This is where that idea ended up.

prtOrd is an external program, ORD500, and the source of ORD500 is not in the codebase. Its printer file is (we saw it among the 12 references), but the program is not. This happens all the time: some code was lost, some is out of scope. The tool works with what we have and shows us what is missing. Here we have a question for the client.

Notes: collecting decisions

Until now we were “just” understanding, but we do that for a purpose, which is to make decisions.

We capture those decisions through notes. We can right-click on a file or a procedure and add a note.

Notes have types. We can use them to say we want to:

  • drop something
  • merge something into something else
  • rename something
  • clarify something, because it is unclear
  • or just leave a comment

For example:

  • ORD500 is missing: we add a note to ask the client about it.
  • The name S01ACT is not exactly self-explanatory: we suggest a rename.
  • CloseCUSTOME1 seems to be unused: we mark it for removal.
  • DAT001 and DAT002 could perhaps be merged, because DAT002 covers just a corner case.

When we are done, we save everything in a single file. It contains the analysis data together with all our notes. That is the basis for our plan.

Reports

Navigating the code inside the tool is useful, but not everybody will use the tool. So we can generate three reports to share with everybody:

  1. Codebase statistics
  2. Architecture assessment
  3. Change plan

The change plan is the most interesting one for our example. It contains all our notes. If we put a note on a procedure, the report also includes the sequence diagram of that procedure, so the reader gets some context.

Why it works for us

This tool proved very valuable for us. We are always in the position of having to become familiar with a codebase very, very quickly.

It also helps us navigate the code together with the client’s people. Different stakeholders know different parts of the system. Some are technical, some are not. With the tool we can look at the same screen, the same program, the same diagram, and discuss it. And we collect decisions as we take them, so at the end they are exported into a well-organized plan, instead of being lost in meeting notes.

We use it mainly to plan migrations. But the same features are useful to plan a refactoring, or simply to maintain a codebase and get familiar with it.

Where I think this is going

When we started building this tool, I thought of it as something that we needed because we had a few weeks to learn about a codebase developed in decades.

But I am starting to think that tools like this one will be needed by many more software engineering teams in the future.

The fact is that codebases are evolving much faster (I wonder who is the culprit of that…). And this means that it is getting A PAIN to keep up with what is going on. So either we give up or we find a way to maintain an overview of the codebases. If we don’t, we keep accumulating features that are no longer relevant, we get architecture drift, we get repetitions. And we just do not know what we have.

I think this will happen more and more. So we will need tools to see what is in our codebases. In this way we will be able to do one thing: simplify. Otherwise this complexity will blow up in our faces.

One letter when there is something worth sending.

Written by Federico · Torino

Scroll to Top