Khoros Communities moves to The Cloud
Photo by Alex Machado on Unsplash Nearly two years ago, the Khoros Communities Engineering organization began to move from co-located data centers to AWS cloud services. It's been a long process, but we're nearly complete. Technical Manager Aditya Pandurangi AdityaP , was kind enough to discuss the project with me. Q: Can you summarize what cloud services are and how Khoros Communities uses them? A: A cloud service, at a basic level, is a collection of many data centers that are spread across multiple locations and owned by a cloud service provider such as Amazon Web Services (AWS) or Microsoft Azure. The provider offers companies like Khoros a cloud-based platform, infrastructure, and storage services. In this case, the infrastructure to host and manage Khoros Communities. The cloud service provides a user interface or a set of APIs to request and manage specific hardware with the software of our choosing. The cloud service provider handles the hardware and maintenance, and we're in control of how we allocate resources. Before cloud services, companies had to have their own hardware located in a data center (either on-premise or in a colocated data center with multiple companies). This meant that companies managed the acquisition and maintenance of their own hardware in addition to managing network traffic and other management tasks. Handling hardware failures, hardware replacements, and physically moving the hardware was all in the company's purview. Khoros Communities, for example, had our own servers and equipment running in two data centers: one in Europe (Amsterdam) and one in the West Coast US (San Jose, California). In the cloud, we no longer worry about sourcing hardware, moving it physically, dealing with failures and outages, sourcing data center facilities, or paying data center storage rates. We don’t need to have Khoros employees add, remove, troubleshoot, and replace physical machines when we need more resources or if something breaks. In The Cloud™, everything is done for you. Q: Any downsides? A: There will always be sporadic hardware failures and restarts in the cloud that might temporarily make our services unavailable. That said, Khoros has built-in redundancy to handle failures as gracefully as possible. Also, the costs of using a cloud service over the long term are probably a bit higher because we don’t own the hardware and can’t amortize the cost over the duration of ownership. Despite those points, moving to the cloud is very much a net benefit. It’s allowed us to do cool, new things that we couldn't before, as well as offer our customers a better experience. Q: What does this migration mean for Khoros Communities customers? A. This migration enables us to provide flexibility to our customers in ways we couldn't before. We can now scale our resources to the needs of any customer. For example, some customers host large events where they see a 2-3x increase in traffic on already busy communities. We're talking lots of views and demands on performance. If a game company has a huge game release taking place or if a software company is holding a major event, we can make sure we have enough hardware ready and on standby as needed. Behind the scenes, our Engineering organization has much more flexibility, which benefits the customer. Unlike in the data center, we’re able to adjust our resources on the fly. More app servers needed? Give TechOps a couple of minutes to an hour -- done! Our database is getting overloaded and needs to be doubled in size temporarily? Once again, shine the TechOps signal -- done! We’ve been able to handle outages like never before, and we’ve been able to support many more customers and resources. Q: What does our AWS infrastructure look like? A: This is a high-level diagram. Q: What were some of the challenges you faced? A: The new environment presented a few challenges. AWS uses a paradigm of separate, siloed regions for different environments (such as QA vs. Production). This was a change from our datacenter infrastructure that used region-agnostic services. In the cloud, each region requires its own service deployment. This created extra work for teams that migrated their services to the cloud, but it was worth the trouble. Our infrastructure is more resilient -- there isn’t a single shared point of failure. An issue in one region does not affect the others. The migration required several teams within the Community Engineering organization to improve services and update workflows. We broke large, multi-purpose services into smaller parts with a more narrow focus. While that led to improved security, technology updates, and better performance, we had to shake off old habits and adjust to multi-step processes. Q: How big was this project, and how long has it been in progress? A: After a couple of false starts, we created the first JIRA ticket for this project on May 15th, 2018, so I think we can consider that the beginning of the project. This migration has been a massive undertaking that’s taken the effort of many teams. A multi-year project takes endurance. For fun, we hung a child's growth chart (like something you'd use to track height over time) in the San Francisco office. Each week, we'd update the chart. This worked great, until Covid when we shut down the office to shelter in place. We got a chance to go back to the office in April to pick up personal items. The chart was still there and we updated our progress. (We got a little lazy with that chunk in the middle where we just drew a simple line.) Later on, Jake Rozin (JakeRo), an engineer in our Customer Operations group, built a UI for us to track the status of community migration. Here's a simplified, sanitized version of the page as we were just finishing moving communities out of the Amsterdam data center. Q: What kind of coordination did the migration require? How many different teams had to work together to make this happen? A: This effort has taken the coordination of many different groups. The primary teams involved have been Technical Operations (TechOps), Application Operations (ApOps), Information Security (InfoSec), and Rocket. Each of these teams has played a vital daily role in the AWS migration process. In addition, all the teams responsible for individual microservices were involved. These teams have had to convert their services to use our new service deployment pattern in AWS, including Dockerizing their service and coordinating with the groups above to migrate and set up infrastructure. Q: What's the Rocket team? A: Rocket is an Engineering team in the Community organization. Essentially the team acts as the glue between TechOps and Engineering. While TechOps deals with networking and setting up/monitoring our infrastructure, the Rocket team works one layer above. We design and implement deployment patterns and pipelines used to integrate Khoros Communities services with our AWS infrastructure. Engineering teams across the Community organization consult us about using our deployment patterns and the best ways to support different service types. Rocket also builds internal tools. For example, we built one service to determine whether we can scale a community to another node and another to find the correct service to perform cloud-based operations (such as adding/removing targets from a load balancer or creating a CDN distribution). Other projects include: improving our security practices and leveraging new cloud-based security options, such as encrypted parameter stores and policy-based access controls (in conjunction with TechOps and InfoSec) updating our current provisioning and de-provisioning processes to work with the new cloud infrastructure and paradigm (in conjunction with TechOps) Q: Did the Rocket team have to build services or tools for the migration process? A: We did! Unfortunately, we are in the process of patenting them. Once we get those filed, we'll write another post with the details 😁 Q: What have been the major milestones for the project? A: We’ve had several major milestones throughout the AWS migration. Proof of concept: Our first milestone in 2018 was building and testing the AWS environment and developing a migration plan. We started with a basic cutover process. This initial proof of concept proved that Khoros could effectively host communities in AWS. The plan slowly evolved to more detailed cutover steps and enabled us to plan for the complete migration. Atlas (Khoros's community) was among the first communities hosted in the cloud. Iteration and automation: Process iteration led to more detailed cutover steps as well as automation services. That brought us to our next major milestone, where we could pass customer migration tasks to our AppOps team and free up Rocket team resources. EMEA community migration and first data center shutdown: End of October/mid-November of 2019, we moved every EMEA customer out of the Amsterdam colocation and into AWS. From there, TechOps completely shut down the Amsterdam data center. AMER community migration: As of Apr 14th, we reached our next milestone. All our AMER communities are in AWS. All that remains are the final milestones -- completing service migration and shutting down the US datacenter in San Jose. Once this happens at the end of June, our AWS project will finally be complete! Q: Can you share and charts or metrics showing improved performance for customers due to the migration? A: Sure. The following charts show the drop in beacon times pre and post-AWS migration. Beacon time is our way of measuring the time it takes a human (we filter out bots) to load a page on a community. Lower beacon times signal that the page being viewed loaded faster. (The numbers on the Y-axis are time in milliseconds.) Two of the customers featured in these charts have heavily trafficked communities. I picked the third at random. You can see that they all showed nice improvements post-migration. Customer 1 Customer 2 Customer 3 Q: Nice! What is it about AWS that lowers beacon times? A: There are a few factors. Our servers in the datacenter were starting to age and the AWS servers are newer, have better CPUs, and thus perform better. In addition, we started routing all traffic through Cloudfront as a CDN, which serves as a caching layer. I imagine that AWS also optimizes the route by which Cloudfront reaches the load balancers for speed. Thanks so much, Aditya! We appreciate your insight and all the details. We'll end here, but before we do, let's give a shout out to all the folks who made the AWS migration possible: Rocket: JonL,eddielo), AdityaP , LauraPe , KaranS , hernan_vinuesa TechOps: BillKr, CanC, ChrisSa, DanielA, DavidSu, EricV, GeorgeB, GokulN, KunjalS, MarcS, MarkJ, MattW, MichaelCa, TauqeerA, WillY Information Security: BryanM, JuanCo, ManjunathM, MinhN, MohanaC, PeterN, SoumyaR, SuyashM Application Operations: kh-mso, ArunkumarG, DimitarI, Georgi Todorov, MichaelM, NicholasD, RonT, (WeiS)1.6KViews
Sign in to react to this post5Comments
How Maia is Powered by Flow.ai and Modern Chat
Maia is a new virtual assistant that Khoros built to help our customers directly from within Atlas, and throughout Khoros' website(s). It's the culmination of years of research and development around how to best help users adopt and utilize Khoros products. Maia's name is rooted in Greek mythology, where Maia is the daughter of Atlas and the mother of Hermes (the messenger of the gods). Maia was designed from the ground up to reflect Khoros' company values, with a voice and tone crafted by Natalie Houchins, a talented member of our Information Experience team. Maia is powered by Khoros Modern Chat's Automation Framework, and acts to both direct users to the answers they need or to one of our Atlas Guides, our dedicated chat agents. Maia is capable of filing support tickets, checking product status, and much more. So, what went into the Maia project? What type of technologies are powering it, and how can you leverage these solutions to create your own interactive chat experience for your customers? A Discussion with Maia's Creators In order to tell the story of Maia, we talked to Travis Berryhill, the leader of the Maia project, and Anshul Jain, the engineer that led in Maia's development. What inspired the creation of Maia? Travis Berryhill: For the longest time, chatbots just weren’t up to task when it came to actually being helpful. So, we wanted to deliver a chatbot that would actually be helpful. A bot that could complete tasks for users and peers and truly make their lives easier. What are some of the technologies the team utilized in Maia's creation? Travis Berryhill: I like to use Lucidchart to plan out the flows before we get started. We use Flow.ai to actually build the bot and train its Natural Language Processing (NLP). We have utilized several APIs, like Statuspage and Salesforce in order to help Maia accomplish tasks. Maia is powered by an amazing startup called Flow.ai, which happens to have just joined the Khoros family mainly because of the superior bots they were able to produce. Flow’s artificial intelligence and machine learning functionality is incredible, and it gives Maia the potential to be a game changer when it comes to bots. What made Flow.ai the platform of choice for Maia? Travis Berryhill: There are so many reasons! Flow.ai had the cleanest UI out of our choices, so that was what caught me at first glance. Then, once I dug into the functionality, I was blown away. Its artificial intelligence and machine learning capabilities are incredible, and it had Maia quickly discovering intents that weren’t even programmed. Flow is also continually innovating. It feels like it has new functionality all the time, and it all just made our lives easier. From a development standpoint, are there any challenges that you faced when creating Maia? How did you overcome them? Anshul Jain: Initially, there were a couple of challenges to understanding the process and how chatbots work, but Flow.ai's great documentation and the Khoros Community Developer Docs made life easier. From a technical point of view, whenever I find any issue either with the code or in the flows, I just check the documentation. Most of the time, I find my solution. Occasionally, I get issues while calling the APIs, so before using any API, I test that in Postman first and then use it in chatbot. What advice would you give a developer tasked with creating their own interactive chatbot for the Automation Framework? Anshul Jain: I would suggest starting with simple flows. Once you get a mindset on how the chatbot can solve the manual process, then you will start getting more ideas. To develop the chatbot, I think that the developer should be proficient at JavaScript and problem-solving. And before starting any development, write down all the steps. Once you are clear with the steps, then start the development. Is there anything exciting coming to Maia we can talk about? We are about to release a flow that will allow users to open support cases from wherever they are instead of having to go to Atlas, every time. I’m really excited about that. Development Nitty Gritty Maia is powered by Khoros Care and its incredibly versatile Automation Framework. The Automation Framework enables you to connect chatbots to Modern Chat and facilitate direct interactions with customers, as well as to hand off those interactions to live agents when the conversation calls for a human touch. Adding the words and interaction trees to Maia was accomplished using Flow.ai. Flow handles the message responses throughout the entire automated conversation. It also initiates the handoff when and if it's time for a human to step in. The actual connection between the chatbot and the customer is all handled through Khoros Care and its Modern Chat. The Automation Framework API facilitates the exchange between Flow.ai and the customer, and then hands the conversation off to an agent when Flow.ai triggers the handoff. From there, an Agent signed in to Khoros Care can interact directly with the customer in one seamless experience. Maia's Incredible Team No discussion about Maia is complete without recognizing the many folks that helped make Maia possible. In addition to Travis Berryhill (TravisB) and Anshul Jain (AnshulJ), a whole host of people contributed to Maia's continued success. These folks include: Natalie Houchins (NatalieH) Jamila Rowser (JamilaR) Lisa Ingram (LisaI) Josh Snider (JoshSn) Justin Fellers (JustinF) Christopher Black (ChristopherB) Jan Morris (JanM)778Views
Sign in to react to this post1Comment
How we built it: Search Anywhere and the Resource Center
In 2019, the Customer Experience team at Khoros planned to improve the digital experience of our customers through a series of initiatives. One of the primary tasks was to make Atlas the focal point of documentation and support articles for Marketing products, just like it was for Care products. The change involved two steps: Moving all our product technical and functional documentation from third-party software to Atlas Maintain the ability to search for documentation from within the product websites i.e. search for documentation from Marketing and Care platforms from the in-app Resource Center What is the Resource Center? The Resource Center in this post refers to the tool powered by third-party software (Pendo) to house additional contextual help and include standard help articles or FAQs, as well as in-app Guides that will walk users through specific processes. Most of the existing Marketing customers were frequent users of the search within the app tool. So, it was important for us to retain this experience. However, these changes were due in less than two months, given the licensing deadline with a third-party application. This presented two key challenges: How to look up information from the Community? How to embed the Search Anywhere tool (yes, that’s what we named it internally) into the Resource Center and into the product website? How we solved looking up information from Community: Community platform supports the awesome API layer, known in developer circles, as LiQL - the Lithium Query Language LiQL enables you to search Community information based on tags, title etc. However, you must have the “correct” permissions via the API keys in order to pull the information The Triumph team built the oAuth based mechanism to provide the API access from a web application The premise was that we could register the Search Anywhere tool as an app on the Community platform and use the specific keys to make API calls and this is how it works currently A key feature that was necessary was to limit the ability to only lookup information limited to only certain “boards” (or nodes if you happen to know Community well). We achieved this by attaching a unique role, based on the application (Marketing or Care) We registered unique apps for Care and Marketing to “sandbox” access to relevant information within each of the products We needed the ability to show some posts by “default”, for example, if you land on the “Social Marketing” page in the Marketing product, the tool would show “top” posts from the “Social Marketing” board in our Community We achieved this using “tags”, a way to label the posts in Community. This required identifying and tagging the relevant posts in specific boards in the Community. Our Product Content Experience team helped with this laborious but important task. Going the extra mile -- Integrating Resource Center into Community and Community Analytics Following the introduction of the Resource Center embedded into the Marketing product, we wished to integrate Pendo and the Resource Center in Community and Community Analytics in order to provide a consistent help experience for our customers across all of Khoros products. The community Admin, Studio, Moderation Manager, and Toolbox areas have a plethora of options, and community admins are required to understand their use and purpose for configuring community as per their brand’s needs. To improve the customer experience, Pendo integration is done with Community to provide dynamic help on the admin section. We leveraged the Search Anywhere work to identify the right resources for Pendo and created the Community integration. Abhishek Gupta, Vaibhav Chawla, and John Dowden collaborated for integrating and creating Search Anywhere built with Atlas tags associated with hundreds of Atlas articles. Search Anywhere makes use of Tags (not Custom Tags like we have in the legacy Admin/Studio Help drawer) to fetch the right information. The sheer number of Tags required to cover the Admin, Studio, Moderation Manager, API Browser, and Toolbox sections was quite a challenge. To top that up, Community Analytics had to be identified through a single Search Anywhere build. The team put in a lot of effort and created a JSON mapping to make it dynamic and easy to upgrade as well as completing a small but very important task to improve the customer experience. For the technically inclined: Vanilla is the way The tool is written in JavaScript and modularized for easy readability Tokens The tool uses Community’s OAuth APIs to obtain and refresh access tokens Build We dockerized the build to make it easy for incremental updates This also helps in developing and testing the tool on any machine Since different apps were registered, we parameterized the build for the code to “know” the context and use appropriate access keys to invoke APIs. If you build this tool with the parameter “CARE”, the output code can only access information from the “Care” board in our Community Deployment The binaries (word for the files that is output from all the code we write) are pushed to an S3 bucket from where the code is copied into the resource center The project was a unique and proactive collaboration between Engineering, Customer Experience and Product Content Experience teams. It did lay the foundation for great working relationships among those involved in the project, very much embodying our value: We win and grow as one team. Callouts: Santosh Shaastry ( SantoshS) Abhinn Gautam ( AbhinnG ) Kokil Jain ( KokilJ ) Narendra Prabhu ( NarendraG ) Gunaalan ( gunaas ) Keerthana ( KeerthanaS ) Annu A ( AnnuA ) Abhishek Gupta ( AbhishekGu ) Vaibhav Chawla ( VaibhavC ) Akash Navani ( AkashN ) John Dowden ( JohnD ) Scott Scarborough ( ScottSc ) Travis Berryhill ( TravisB )901Views
Sign in to react to this post7Comments
You’ve seen all recent content