Welcome!

Microservices Expo Authors: Aruna Ravichandran, Elizabeth White, Liz McMillan, Pat Romanski, Cameron Van Orman

Related Topics: Microservices Expo, Java IoT, Microsoft Cloud, Machine Learning , Agile Computing, @CloudExpo

Microservices Expo: Article

Can You See the Storm Coming?

APM solutions enables us to set up alerts against good performance baselines

As much as we try to avoid performance problems, they do happen. It is inevitable. But it is possible to learn to react fast, and in some occasions fast enough that the impact on the end users is negligible. Despite operators' best efforts, 73% of performance issues are reported by users, according to "APM: Getting IT on the C-Level's agenda" report by Aberdeen Group. This number is quite large considering that less than 5% of all users typically bother to complain at all. User Experience has a significant impact on business success. According to the Aberdeen report, poor performance of applications can reduce revenue by 9% and productivity by 64%.

The goal of application performance monitoring is to ensure and improve the quality of applications as perceived by the end users. Getting to the root of the problem quickly is only part of the solution. When we ask various Operation teams how they learn about performance problems they sometimes reply: "Our users tell us." As we already pointed it out we should not wait for the disaster to happen, but rather take appropriate actions as soon as we see the storm coming.

In this article we recount two incidents that happened to our client, ZinMines, a steel and mining company from Zinariya (names changed for commercial reasons). In both cases the Operations team at ZinMines got notified about the problem well in advance of any user reaction. The team members were able to start analyzing and improving the situation by the time users eventually notified them about the problem. If they have waited for users to notify them, the problem would have been solved much later and users would have been much more frustrated.

Case #1: The Maintenance Page
When the Operations team first set up its application performance monitoring solution, the members made sure that the alerts on potential performance problems were set up correctly.

One morning, just before 8 a.m., an alert that monitored total transaction time went off. The Operations team used the APM solution to chart the total time as seen by the end user broken down by time spent on the server and time spent on the network. They saw that significant time was spent on the server (see Figure 1). This could make the whole application slower.

The team started to analyze the problem together with the engineering team. They learned that there was a serious bug and the engineering team would need a few hours to fix it.

Meanwhile, shortly after 9 a.m., they got a call from a user that the services run by ZinMines were particularly slow. The helpdesk informed the users that the problem was already being investigated and a team had been appointed to look into the issue.

Since the delays in processing user requests kept coming in, and in order to avoid further frustration among the end users, the decision was made to enter into maintenance mode. Around 9:30 the Operations team started redirecting part of the traffic to the maintenance page. This took some load off the application, decreased total time and gave the engineering team time to handle the issue. The chart in Figure 1 shows the change in server time and redirect time after the redirect to the maintenance page was enabled, which is indicated on the time line with the blue arrow.

Figure 1 also shows that the traffic was already pretty high for at least one hour before one of the end users notified the Operations team about the problem. The red arrow on the time line with the red arrow in Figure 1 indicates when the problem was first reported.

Figure 1: Reduced server time and total time after activating redirect to the maintenance page

Had the Operations team waited for an end user to report the problem they would not have contacted the engineering team early enough to gain extra time to start resolving the problem. Thanks to properly configured alerts they were able to act in time and shorten the time the users were impacted by poor application performance.

Case #2: 4xx Errors
Sometime later, the Operations team got another alert. They consulted the APM solution and discovered that one of the application servers was generating a lot of 404 errors.

Figure 2 shows the 4xx errors charted around the time of the incident report. The green arrow indicates when the alert was raised.

Figure 2: 4xx errors were happening for almost an hour before the incident was reported by the end user and fixed.

The Operations team performed root cause analysis and discovered that the problem was caused by some caching issues and they decided to restart the server. The blue arrow on the time line in Figure 2 shows when the server was restarted.

Shortly before they initiated the restart procedures they got an incident report submitted by one of the end users from the finance department (see Figure 3).

Figure 3: Incident reported an hour after the performance problems had started.

They could close the issue almost immediately since when they checked the report with HTTP 4xx errors (see Figure 2) the situation was back to normal again. This was again seen as a very positive element by the customer. The team was not only aware of the issues that were troubling the users, before any user would complain, but could also see if the action taken to resolve the issue actually improved the situation.

The red arrow in Figure 2 shows when the report was submitted by the end user. The problem started more than an hour before the user reported the incident, similar to the previous incident where users waited more than an hour. If the Operations team waited for their users they would have lost at least an hour resolving the issue.

Conclusion
Application performance monitoring is more than just following the fault domain isolation workflow to determine the root cause of the problem that is reported by an end user. In many cases waiting for the end users to report problems simply takes too long, while the frustration due to poor performance grows.

APM solutions, such as Compuware dynaTrace Data Center Real User Monitoring (DCRUM), enables us to set up alerts, e.g., against good performance baselines. In many cases these alerts will be triggered long before the end users report an incident so it gives the Operations team more time to react before issues get serious. As shown in the Aberdeen report mentioned above 95% of users will not even bother to report the problem at all.

Second it also enables the team to see if the action taken to rectify the problem actually improves the end-users' experience.

More Stories By Sebastian Kruk

Sebastian Kruk is a Technical Product Strategist, Center of Excellence, at Compuware APM Business Unit.

Comments (0)

Share your thoughts on this story.

Add your comment
You must be signed in to add a comment. Sign-in | Register

In accordance with our Comment Policy, we encourage comments that are on topic, relevant and to-the-point. We will remove comments that include profanity, personal attacks, racial slurs, threats of violence, or other inappropriate material that violates our Terms and Conditions, and will block users who make repeated violations. We ask all readers to expect diversity of opinion and to treat one another with dignity and respect.


@MicroservicesExpo Stories
Digital transformation leaders have poured tons of money and effort into coding in recent years. And with good reason. To succeed at digital, you must be able to write great code. You also have to build a strong Agile culture so your coding efforts tightly align with market signals and business outcomes. But if your investments in testing haven’t kept pace with your investments in coding, you’ll lose. But if your investments in testing haven’t kept pace with your investments in coding, you’ll...
In his session at 21st Cloud Expo, Michael Burley, a Senior Business Development Executive in IT Services at NetApp, will describe how NetApp designed a three-year program of work to migrate 25PB of a major telco's enterprise data to a new STaaS platform, and then secured a long-term contract to manage and operate the platform. This significant program blended the best of NetApp’s solutions and services capabilities to enable this telco’s successful adoption of private cloud storage and launchi...
DevOps at Cloud Expo, taking place October 31 - November 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA, is co-located with 21st Cloud Expo and will feature technical sessions from a rock star conference faculty and the leading industry players in the world. The widespread success of cloud computing is driving the DevOps revolution in enterprise IT. Now as never before, development teams must communicate and collaborate in a dynamic, 24/7/365 environment. There is no time to w...
Enterprises are adopting Kubernetes to accelerate the development and the delivery of cloud-native applications. However, sharing a Kubernetes cluster between members of the same team can be challenging. And, sharing clusters across multiple teams is even harder. Kubernetes offers several constructs to help implement segmentation and isolation. However, these primitives can be complex to understand and apply. As a result, it’s becoming common for enterprises to end up with several clusters. Thi...
Containers are rapidly finding their way into enterprise data centers, but change is difficult. How do enterprises transform their architecture with technologies like containers without losing the reliable components of their current solutions? In his session at @DevOpsSummit at 21st Cloud Expo, Tony Campbell, Director, Educational Services at CoreOS, will explore the challenges organizations are facing today as they move to containers and go over how Kubernetes applications can deploy with lega...
Today most companies are adopting or evaluating container technology - Docker in particular - to speed up application deployment, drive down cost, ease management and make application delivery more flexible overall. As with most new architectures, this dream takes significant work to become a reality. Even when you do get your application componentized enough and packaged properly, there are still challenges for DevOps teams to making the shift to continuous delivery and achieving that reducti...
Is advanced scheduling in Kubernetes achievable? Yes, however, how do you properly accommodate every real-life scenario that a Kubernetes user might encounter? How do you leverage advanced scheduling techniques to shape and describe each scenario in easy-to-use rules and configurations? In his session at @DevOpsSummit at 21st Cloud Expo, Oleg Chunikhin, CTO at Kublr, will answer these questions and demonstrate techniques for implementing advanced scheduling. For example, using spot instances ...
SYS-CON Events announced today that Cloud Academy has been named “Bronze Sponsor” of SYS-CON's 21st International Cloud Expo®, which will take place on Oct. 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Cloud Academy is the leading technology training platform for enterprise multi-cloud infrastructure. Cloud Academy is trusted by leading companies to deliver continuous learning solutions across Amazon Web Services, Microsoft Azure, Google Cloud Platform, and the most...
The last two years has seen discussions about cloud computing evolve from the public / private / hybrid split to the reality that most enterprises will be creating a complex, multi-cloud strategy. Companies are wary of committing all of their resources to a single cloud, and instead are choosing to spread the risk – and the benefits – of cloud computing across multiple providers and internal infrastructures, as they follow their business needs. Will this approach be successful? How large is the ...
DevOps is often described as a combination of technology and culture. Without both, DevOps isn't complete. However, applying the culture to outdated technology is a recipe for disaster; as response times grow and connections between teams are delayed by technology, the culture will die. A Nutanix Enterprise Cloud has many benefits that provide the needed base for a true DevOps paradigm. In their Day 3 Keynote at 20th Cloud Expo, Chris Brown, a Solutions Marketing Manager at Nutanix, and Mark Lav...
Many organizations adopt DevOps to reduce cycle times and deliver software faster; some take on DevOps to drive higher quality and better end-user experience; others look to DevOps for a clearer line-of-sight to customers to drive better business impacts. In truth, these three foundations go together. In this power panel at @DevOpsSummit 21st Cloud Expo, moderated by DevOps Conference Co-Chair Andi Mann, industry experts will discuss how leading organizations build application success from all...
DevSecOps – a trend around transformation in process, people and technology – is about breaking down silos and waste along the software development lifecycle and using agile methodologies, automation and insights to help get apps to market faster. This leads to higher quality apps, greater trust in organizations, less organizational friction, and ultimately a five-star customer experience. These apps are the new competitive currency in this digital economy and they’re powered by data. Without ...
With the modern notion of digital transformation, enterprises are chipping away at the fundamental organizational and operational structures that have been with us since the nineteenth century or earlier. One remarkable casualty: the business process. Business processes have become so ingrained in how we envision large organizations operating and the roles people play within them that relegating them to the scrap heap is almost unimaginable, and unquestionably transformative. In the Digital ...
These days, APIs have become an integral part of the digital transformation journey for all enterprises. Every digital innovation story is connected to APIs . But have you ever pondered over to know what are the source of these APIs? Let me explain - APIs sources can be varied, internal or external, solving different purposes, but mostly categorized into the following two categories. Data lakes is a term used to represent disconnected but relevant data that are used by various business units wit...
The nature of the technology business is forward-thinking. It focuses on the future and what’s coming next. Innovations and creativity in our world of software development strive to improve the status quo and increase customer satisfaction through speed and increased connectivity. Yet, while it's exciting to see enterprises embrace new ways of thinking and advance their processes with cutting edge technology, it rarely happens rapidly or even simultaneously across all industries.
With the rise of DevOps, containers are at the brink of becoming a pervasive technology in Enterprise IT to accelerate application delivery for the business. When it comes to adopting containers in the enterprise, security is the highest adoption barrier. Is your organization ready to address the security risks with containers for your DevOps environment? In his session at @DevOpsSummit at 21st Cloud Expo, Chris Van Tuin, Chief Technologist, NA West at Red Hat, will discuss: The top security r...
Most of the time there is a lot of work involved to move to the cloud, and most of that isn't really related to AWS or Azure or Google Cloud. Before we talk about public cloud vendors and DevOps tools, there are usually several technical and non-technical challenges that are connected to it and that every company needs to solve to move to the cloud. In his session at 21st Cloud Expo, Stefano Bellasio, CEO and founder of Cloud Academy Inc., will discuss what the tools, disciplines, and cultural...
Enterprises are moving to the cloud faster than most of us in security expected. CIOs are going from 0 to 100 in cloud adoption and leaving security teams in the dust. Once cloud is part of an enterprise stack, it’s unclear who has responsibility for the protection of applications, services, and data. When cloud breaches occur, whether active compromise or a publicly accessible database, the blame must fall on both service providers and users. In his session at 21st Cloud Expo, Ben Johnson, C...
21st International Cloud Expo, taking place October 31 - November 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA, will feature technical sessions from a rock star conference faculty and the leading industry players in the world. Cloud computing is now being embraced by a majority of enterprises of all sizes. Yesterday's debate about public vs. private has transformed into the reality of hybrid cloud: a recent survey shows that 74% of enterprises have a hybrid cloud strategy. Me...
‘Trend’ is a pretty common business term, but its definition tends to vary by industry. In performance monitoring, trend, or trend shift, is a key metric that is used to indicate change. Change is inevitable. Today’s websites must frequently update and change to keep up with competition and attract new users, but such changes can have a negative impact on the user experience if not managed properly. The dynamic nature of the Internet makes it necessary to constantly monitor different metrics. O...