ASUG News + Views
Dis­as­ter Recov­ery and Busi­ness Con­ti­nu­ity Plan Saved $40 Million
Donald Hook May 2, 2020
Bookmark
Share Article:

A food and bev­er­age dis­tri­b­u­tion com­pa­ny had aggres­sive plans to expand its busi­ness by 40% over four years. In order to achieve this goal, the com­pa­ny rec­og­nized that it need­ed to mod­ern­ize and posi­tion its IT to sup­port this growth. 

This wasn’t an easy task as IT his­tor­i­cal­ly was viewed as a cost cen­ter. Exec­u­tives had not invest­ed in infra­struc­ture or solu­tions for more than 12 years. The result was a very dat­ed envi­ron­ment with appli­ca­tions and infra­struc­ture no longer eli­gi­ble for main­te­nance or support. 

To help move the project in the right direc­tion, the com­pa­ny brought me in as an IT leader to define a strat­e­gy and road map for the next three years. In this arti­cle, I’ll explain the steps we took to cre­ate the strat­e­gy and road map, along with some lessons learned along the way. 

Defin­ing and Sell­ing the IT Strat­e­gy and Road Map

This was the first time this com­pa­ny put in place an IT strat­e­gy and road map, so it was dif­fi­cult to get cor­po­rate buy-in. It was impor­tant to include the C‑suite and busi­ness lead­ers in the plan­ning process to make sure we had a good under­stand­ing of what was impor­tant and how to move forward. 

The first ini­tia­tive we iden­ti­fied was to imple­ment dis­as­ter recov­ery (DR) and busi­ness con­ti­nu­ity plan­ning (BCP). The com­pa­ny was also work­ing on an SAP upgrade from SAP ECC 5.0 to SAP ECC 6.0 SP7 on an IBM iSeries. Although there were dai­ly tape back­ups, there was noth­ing to restore the back­up to. The com­pa­ny had no plan in place to con­tin­ue to oper­ate if it lost the main data cen­ter and its SAP appli­ca­tions. This was a huge risk. 

We had to work to explain the risk and the need for the dis­as­ter recov­ery and busi­ness con­ti­nu­ity plan. Part of the process was explain­ing the finan­cial impact that such a risk can have on the com­pa­ny. Even­tu­al­ly, we did get buy-in and began to imple­ment the plan. 

From Build­ing to Imple­ment­ing the Dis­as­ter Recov­ery and Busi­ness Con­ti­nu­ity Plan

The IT ini­tia­tive would first address the dis­as­ter recov­ery plan and then the busi­ness con­ti­nu­ity plan. 

The dis­as­ter recov­ery por­tion was more tech­ni­cal in nature. We worked with our SAP part­ner to imple­ment a back­up site for our SAP instances. This involved deter­min­ing the size and capac­i­ty of the remote instance and mak­ing sure it would have the capac­i­ty to sup­port the remote offices if we need­ed to fail over to it. 

We also need­ed to increase the band­width from our pri­ma­ry data cen­ter to our partner’s data cen­ter. This was nec­es­sary because we had a repli­ca­tion agent run­ning on the iSeries that sent SAP updates to the remote instance. The non-SAP appli­ca­tions crit­i­cal to run­ning the busi­ness were vir­tu­al­ized. We need­ed con­fir­ma­tion that our vir­tu­al­ized envi­ron­ment was being repli­cat­ed to our back­up data cen­ter and that the back­ups would work if we need­ed to use them.

The busi­ness con­ti­nu­ity por­tion of the plan required more strat­e­gy. The VP of oper­a­tions was involved to help engage and align the busi­ness units and cor­po­rate func­tions, includ­ing HR and finance, among oth­ers. The IT depart­ment and the busi­ness units worked togeth­er to define the pro­ce­dures and steps to exe­cute if a cat­a­stroph­ic event were to occur. We then com­mu­ni­cat­ed these pro­ce­dures to the busi­ness, which includ­ed defin­ing key roles, respon­si­bil­i­ties, and alter­nate facil­i­ties where we would run dur­ing a disaster. 

With all the pieces in place, the last phase focused on test­ing. We need­ed to know we were pre­pared to exe­cute. This required sev­er­al tests where we stopped the sys­tem to repli­cate an out­age and failed over to our back­up instance. We val­i­dat­ed our sys­tems by run­ning reports and going online to check things that were defined in our test plans, such as orders and finan­cials. Once we were suc­cess­ful in val­i­dat­ing our results, we imple­ment­ed the solution. 

Putting the Busi­ness Con­ti­nu­ity Plan to the Test in Real Time 

With a dat­ed infra­struc­ture envi­ron­ment, one item among many that need­ed to be updat­ed was the back­up bat­tery sys­tem for the data cen­ter. The sys­tem was updat­ed, and every­thing seemed to oper­ate as nor­mal. The next day, how­ev­er, my net­work admin­is­tra­tors report­ed a fire in the data center. 

We need­ed to shut things down but could not get to the elec­tri­cal pan­el in the back of the room. We decid­ed to close the doors and evac­u­ate. We imme­di­ate­ly imple­ment­ed the busi­ness con­ti­nu­ity plan we had just final­ized and dis­patched our sales teams to an alter­nate loca­tion. The next step was to imple­ment the dis­as­ter recov­ery plan. With the help of our SAP part­ner, we were able to exe­cute the plan right away. We exe­cut­ed the dis­as­ter recov­ery and busi­ness con­ti­nu­ity plan with­in 20 minutes.

After assess­ing the dam­age, we found the data cen­ter was unre­cov­er­able. We lost the servers, switch­es, iSeries, tele­phones, bat­tery back­up, and the inven­to­ry of lap­tops and parts. We assem­bled the IT team and start­ed to bring up the non-SAP appli­ca­tions from our back­up facil­i­ty. The IT team val­i­dat­ed the con­nec­tiv­i­ty between appli­ca­tions, SAP, and the var­i­ous data­bas­es. The sales teams were already at the remote office loca­tion resum­ing order entry and pro­cess­ing orders. The only issue we expe­ri­enced was lost access to a non-crit­i­cal appli­ca­tion due to a recent pass­word change in the active direc­to­ry. We con­tin­ued to mon­i­tor all the sys­tems to make sure every­thing was run­ning prop­er­ly. Lat­er that evening, the team con­firmed that all orders and deliv­er­ies were processed on our nor­mal schedule. 

Resum­ing Busi­ness as Nor­mal Post-Disaster

Nev­er in a mil­lion years would I ever have expect­ed to encounter a dis­as­ter that would require us to imple­ment our con­tin­gency plan. Because of the work we did to plan, the com­pa­ny con­tin­ued to oper­ate on our back­up facil­i­ties for 40 days. We need­ed to restore our net­work infra­struc­ture, servers, and the iSeries in our main data center. 

The com­pa­ny did rough­ly $1 mil­lion in orders per day. We were able to save the com­pa­ny a sig­nif­i­cant amount of mon­ey by hav­ing our Dis­as­ter Recov­ery and Busi­ness Con­ti­nu­ity Plan in place. Some­thing that is prob­a­bly more impor­tant than mon­ey is the company’s rep­u­ta­tion with its cus­tomers. We were able to con­tin­ue to take care of our cus­tomers as if noth­ing had ever hap­pened. Quite an accom­plish­ment, with all things con­sid­ered — all thanks to hav­ing a plan in place. 

Learn more on how to respond to uncer­tain­ty in CIOs Dis­cuss How to Nav­i­gate the Uncer­tain­ty of COVID-19.” We wel­come all ASUG mem­bers to sub­mit their ideas for blogs they would like to write. 

You Might Be Interested In


Insights Included in Membership
View All Insights
Bookmark
Bookmark
Bookmark
Bookmark