Team Ai
Apppublic

freeEDU/Log-Decoder

sourceHugging Facemitupdated 3y agoView on Hugging Face
0likes
app.py315 linesDownload Raw Back to data
1# To run streamlit, go to terminal and type: 'streamlit run app.py'2# Core Packages ###########################3import streamlit as st4import os5 6import time7import numpy as np8import pandas as pd9 10from loglizer.models import PCA11from loglizer import dataloader, preprocessing12import openai13 14openai.api_key = os.environ["FREEEDU_OPENAI_API_KEY"]15 16project_title = "ChatGPT Log Decoder"17project_desc = """18Intrusion Detection Systems (IDS) are powerful security tools that monitor network traffic for suspicious activity and issues alerts when such activities are discovered. 19While these systems are commonly used by cybersecurity professionals and IT experts, 20the technical jargon and analysis that an IDS provides can be incomprehensible to the average user. 21This project aims to bridge this gap by utilizing ChatGPT's advanced natural language processing to 22articulate the technical information received from an IDS into human-readable form, allowing common everyday users 23to grasp the extent of their network's security status without the need for a technical background.24"""25 26project_link = "https://github.com/logpai/loglizer"27project_icon = "46_Heatmap.png"28st.set_page_config(page_title=project_title, initial_sidebar_state='collapsed', page_icon=project_icon)29 30# additional info from the readme31add_info_md = """32 33# loglizer34 35 36**Loglizer is a machine learning-based log analysis toolkit for automated anomaly detection**. 37  38 39Logs are imperative in the development and maintenance process of many software systems. They record detailed40runtime information during system operation that allows developers and support engineers to monitor their systems and track abnormal behaviors and errors. Loglizer provides a toolkit that implements a number of machine-learning based log analysis techniques for automated anomaly detection. 41 42:telescope: If you use loglizer in your research for publication, please kindly cite the following paper.43+ Shilin He, Jieming Zhu, Pinjia He, Michael R. Lyu. [Experience Report: System Log Analysis for Anomaly Detection](https://jiemingzhu.github.io/pub/slhe_issre2016.pdf), *IEEE International Symposium on Software Reliability Engineering (ISSRE)*, 2016. [[Bibtex](https://dblp.org/rec/bibtex/conf/issre/HeZHL16)][[中文版本](https://github.com/AmateurEvents/article/issues/2)]44**(ISSRE Most Influential Paper)**45 46## Framework47 48![Framework of Anomaly Detection](/docs/img/framework.png)49 50The log analysis framework for anomaly detection usually comprises the following components:51 521. **Log collection:** Logs are generated at runtime and aggregated into a centralized place with a data streaming pipeline, such as Flume and Kafka. 532. **Log parsing:** The goal of log parsing is to convert unstructured log messages into a map of structured events, based on which sophisticated machine learning models can be applied. The details of log parsing can be found at [our logparser project](https://github.com/logpai/logparser).543. **Feature extraction:** Structured logs can be sliced into short log sequences through interval window, sliding window, or session window. Then, feature extraction is performed to vectorize each log sequence, for example, using an event counting vector. 554. **Anomaly detection:** Anomaly detection models are trained to check whether a given feature vector is an anomaly or not.56 57 58## Models59 60Anomaly detection models currently available:61 62| Model | Paper reference |63| :--- | :--- |64| **Supervised models** |65| LR | [**EuroSys'10**] [Fingerprinting the Datacenter: Automated Classification of Performance Crises](https://www.microsoft.com/en-us/research/wp-content/uploads/2009/07/hiLighter.pdf), by Peter Bodík, Moises Goldszmidt, Armando Fox, Hans Andersen. [**Microsoft**] |66| Decision Tree | [**ICAC'04**] [Failure Diagnosis Using Decision Trees](http://www.cs.berkeley.edu/~brewer/papers/icac2004_chen_diagnosis.pdf), by Mike Chen, Alice X. Zheng, Jim Lloyd, Michael I. Jordan, Eric Brewer. [**eBay**] |67| SVM | [**ICDM'07**] [Failure Prediction in IBM BlueGene/L Event Logs](https://www.researchgate.net/publication/4324148_Failure_Prediction_in_IBM_BlueGeneL_Event_Logs), by Yinglung Liang, Yanyong Zhang, Hui Xiong, Ramendra Sahoo. [**IBM**]|68| **Unsupervised models** |69| LOF | [**SIGMOD'00**] [LOF: Identifying Density-Based Local Outliers](), by Markus M. Breunig, Hans-Peter Kriegel, Raymond T. Ng, Jörg Sander. |70| One-Class SVM | [**Neural Computation'01**] [Estimating the Support of a High-Dimensional Distribution](), by John Platt, Bernhard Schölkopf, John Shawe-Taylor, Alex J. Smola, Robert C. Williamson. |71| Isolation Forest | [**ICDM'08**] [Isolation Forest](https://cs.nju.edu.cn/zhouzh/zhouzh.files/publication/icdm08b.pdf), by Fei Tony Liu, Kai Ming Ting, Zhi-Hua Zhou. |72| PCA | [**SOSP'09**] [Large-Scale System Problems Detection by Mining Console Logs](http://iiis.tsinghua.edu.cn/~weixu/files/sosp09.pdf), by Wei Xu, Ling Huang, Armando Fox, David Patterson, Michael I. Jordan. [**Intel**] |73| Invariants Mining | [**ATC'10**] [Mining Invariants from Console Logs for System Problem Detection](https://www.usenix.org/legacy/event/atc10/tech/full_papers/Lou.pdf), by Jian-Guang Lou, Qiang Fu, Shengqi Yang, Ye Xu, Jiang Li. [**Microsoft**]|74| Clustering | [**ICSE'16**] [Log Clustering based Problem Identification for Online Service Systems](https://www.microsoft.com/en-us/research/wp-content/uploads/2016/07/ICSE-2016-2-Log-Clustering-based-Problem-Identification-for-Online-Service-Systems.pdf), by Qingwei Lin, Hongyu Zhang, Jian-Guang Lou, Yu Zhang, Xuewei Chen. [**Microsoft**]|75| DeepLog (coming)| [**CCS'17**] [DeepLog: Anomaly Detection and Diagnosis from System Logs through Deep Learning](https://www.cs.utah.edu/~lifeifei/papers/deeplog.pdf), by Min Du, Feifei Li, Guineng Zheng, Vivek Srikumar. |76| AutoEncoder (coming)| [**Arxiv'18**] [Anomaly Detection using Autoencoders in High Performance Computing Systems](https://arxiv.org/abs/1811.05269), by Andrea Borghesi, Andrea Bartolini, Michele Lombardi, Michela Milano, Luca Benini. |77 78 79## Log data80We have collected a set of labeled log datasets in [loghub](https://github.com/logpai/loghub) for research purposes. If you are interested in the datasets, please follow the link to submit your access request.81 82## Install83```bash84git clone https://github.com/logpai/loglizer.git85cd loglizer86pip install -r requirements.txt87```88 89## API usage90 91```python92# Load HDFS dataset. If you would like to try your own log, you need to rewrite the load function.93(x_train, y_train), (x_test, y_test) = dataloader.load_HDFS(...)94 95# Feature extraction and transformation96feature_extractor = preprocessing.FeatureExtractor()97feature_extractor.fit_transform(...) 98 99# Model training100model = PCA()101model.fit(...)102 103# Feature transform after fitting104x_test = feature_extractor.transform(...)105# Model evaluation with labeled data106model.evaluate(...)107 108# Anomaly prediction109x_test = feature_extractor.transform(...)110model.predict(...) # predict anomalies on given data111```112 113For more details, please follow [the demo](./docs/demo.md) in the docs to get started. Please note that all ML models are not magic, you need to figure out how to tune the parameters in order to make them work on your own data. 114 115## Benchmarking results 116 117If you would like to reproduce the following results, please run [benchmarks/HDFS_bechmark.py](./benchmarks/HDFS_bechmark.py) on the full HDFS dataset (HDFS100k is for demo only).118 119|       |            | HDFS |     |120| :----:|:----:|:----:|:----:|121| **Model** | **Precision** | **Recall** | **F1** |122| LR| 0.955 |	0.911 |	0.933 |123| Decision Tree | 0.998 |	0.998 |	0.998 |124| SVM| 0.959 |	0.970 |	0.965 |125| LOF | 0.967 | 0.561 | 0.710 |126| One-Class SVM | 0.995 | 0.222| 0.363 |127| Isolation Forest |  0.830 | 0.776 | 0.802 |128| PCA | 0.975 | 0.635 | 0.769|129| Invariants Mining | 0.888 | 0.945 | 0.915|130| Clustering | 1.000 | 0.720 | 0.837 |131 132## Contributors133+ [Shilin He](https://shilinhe.github.io), The Chinese University of Hong Kong134+ [Jieming Zhu](https://jiemingzhu.github.io), The Chinese University of Hong Kong, currently at Huawei Noah's Ark Lab135+ [Pinjia He](https://pinjiahe.github.io/), The Chinese University of Hong Kong, currently at ETH Zurich136 137 138## Feedback139For any questions or feedback, please post to [the issue page](https://github.com/logpai/loglizer/issues/new). 140 141"""142#######################################################################################################################143default_log_path = os.path.join("data","HDFS","HDFS_100k.log_structured.csv")144def load_loglizer(train_log_path=default_log_path):145    print("Loading Loglizer PCA model:")146    start = time.time()147    (x_train, _), (_, _), _ = dataloader.load_HDFS(train_log_path, window='session',148                                                split_type='sequential', save_csv=True)149    feature_extractor = preprocessing.FeatureExtractor()150    x_train = feature_extractor.fit_transform(x_train, term_weighting='tf-idf',151                                              normalization='zero-mean')152    model = PCA()153    model.fit(x_train)154 155    stop = time.time()156    print("Loading complete. Time elapsed: " + str(stop - start))157 158    # Extract the event labels for each unique eventid in the source structured log file159    # Load the CSV file160    df = pd.read_csv(train_log_path)161 162    # Extract unique 'EventId' and 'EventTemplate' pairs163    event_names = df[['EventId', 'EventTemplate']].drop_duplicates()164    event_names_dict = dict(zip(event_names['EventId'], event_names['EventTemplate']))165 166    return model, feature_extractor, event_names_dict167 168def analyze_with_chatgpt(anomalous_packets):169    context = {"role":"system",170               "content":"""171               You will receive 5 rows of transformed data logs, where the middlemost log is flagged as anomalous.172               Your job is to analyze and compare this log with its surrounding context logs, and find out why it was flagged as anomalous. 173               Your response should be 50 to 150 words only; be as concise as possible and address the user directly. 174               Only mention the key issues and cause of the anomaly. Never mention the Column Names, the middlemost log, surrounding context log.175               You can mention what the irregular values in a column represent, and base your report on that.176               Assume that the user has no knowledge of networks or cybersecurity. Your response should start with177                "Based on the logs,".178               Try to avoid telling the user 'if the issue persists'.179               Give simple step-by-step advice on what the user can do on their own.180               If the advice is beyond the scope of the everyday user, notify them that the assistance of a professional 181               is necessary for the level of intrusion you found.182               """}183    message_prompt = []184    message_prompt.append(context)185    anomalies = f"*Column Names*: {anomalous_packets.columns} --- "186    # Build a sequence of messages, with each message being a single packet up to the last packet flagged as anomalous187    for i,p in anomalous_packets.iterrows():188        anomalies += f"*Row {i+1}*: "189        anomalies += str(p.values)190        anomalies += " --- "191 192    message_format = {"role":"user",193                      "content":anomalies}194    message_prompt.append(message_format)195 196    # Generate the response197    response = openai.ChatCompletion.create(198        model="gpt-3.5-turbo",199        messages=message_prompt,200        max_tokens=1008,201        temperature=0.7,202        n=1,203        stop=None,204    )205 206    # Extract the response text from the API response207    response_text = response['choices'][0]['message']['content']208    # Return the response text and updated chat history209    return response_text210 211def main():212    head_col = st.columns([1,8])213    with head_col[0]:214        st.image(project_icon)215    with head_col[1]:216        st.title(project_title)217 218    st.write(project_desc)219    st.write(f"Source Project: {project_link}")220    expander = st.expander("Additional Information on the Source Project (Loglizer)")221    expander.markdown(add_info_md)222    st.markdown("***")223    st.subheader("")224#########################################225 226    # instructions and file upload button227    st.subheader("""228    How to use: 229     1. Upload your log file (must be .csv format)230     2. Click the 'Run Loglizer' button231                 """)232    uploaded_file = st.file_uploader("Upload a log file in csv format:",233                               type=["csv"],234                               accept_multiple_files=False)235 236#########################################237 238    run_button = st.button("Run Loglizer")239 240    if "anomalies" not in st.session_state:241        st.session_state.anomalies = []242    if "col_names" not in st.session_state:243        st.session_state.col_names = []244 245    # button is clicked246    if run_button:247        if uploaded_file is None: # if no files were uploaded:248            st.error("Please upload a .csv log file")249 250        else:251            # Ensure temp directory exists252            if not os.path.exists('temp'):253                os.makedirs('temp')254 255            filename = os.path.basename(uploaded_file.name)256            file_path = os.path.join("temp", filename)257            # Write out the uploaded file to temp directory258            with open(file_path, 'wb') as f:259                f.write(uploaded_file.getbuffer())260 261            with st.spinner('Loading Loglizer...'):262                model, feature_extractor, event_names_dict = load_loglizer()263 264            # Simulate the 'live' log feed265            (x_live, _), (_, _), _ = dataloader.load_HDFS(file_path, window='session', split_type='sequential')266            x_live, st.session_state.col_names = feature_extractor.transform(x_live)267            # st.write(f"x_live shape: {x_live.shape}")268 269            # Map event IDs to event templates270            st.session_state.col_names = [event_names_dict[id] if id in event_names_dict else id for id in271                                          st.session_state.col_names]272 273            for i in range(len(x_live)):274                log = x_live[i]275                log = np.expand_dims(log, axis=0)276                prediction = model.predict(log)277                print(f"prediction for x_live[{i}]: {prediction}")278                if prediction == 1:279                    # Anomaly detected280                    context = x_live[max(0, i - 2):min(i + 3, len(x_live))]  # Get the two logs before and after281                    st.session_state.anomalies.append(context)282            # st.write("predict on x_live:")283            # st.write(model.predict(x_live))284 285    st.write("Anomalous events")286    st.write(len(st.session_state.anomalies))287    #288    # st.write("Column names")289    # st.write(st.session_state.col_names)290 291    if len(st.session_state.anomalies) > 0:292        # Displaying the anomalous packets and sending to ChatGPT293        selected_index = st.selectbox('Select an anomaly', list(range(len(st.session_state.anomalies))), 0)294        selected_anomaly = st.session_state.anomalies[selected_index]295        st.write(f"Loglizer found and compiled {selected_anomaly.shape[1]} events from the logs.\n The middle entry has an anomaly:")296        # st.write(f"Shape: {selected_anomaly.shape} Type: {type(selected_anomaly)}")297 298        # Convert the numpy array to a pandas DataFrame299        selected_anomaly_df = pd.DataFrame(selected_anomaly)300        # Set the column names301        selected_anomaly_df.columns = st.session_state.col_names302 303        # st.write(selected_anomaly)304        st.write(selected_anomaly_df)305 306        if st.button('Send to ChatGPT'):307            result = analyze_with_chatgpt(selected_anomaly_df)308            st.write(result)309 310 311if __name__ == '__main__':312    main()313 314# To run streamlit, go to terminal and type: 'streamlit run app-source.py'315