Spaces:

spark-nlp
/

Italian-NER

Sleeping

App Files Files Community

abdullahmubeen10 commited on Jul 24, 2024

Commit

556ef41

verified ·

1 Parent(s): e1b8e3c

Upload 5 files

Browse files

Files changed (5) hide show

.streamlit/config.toml +3 -0
Demo.py +163 -0
Dockerfile +70 -0
pages/Workflow & Model Overview.py +242 -0
requirements.txt +6 -0

.streamlit/config.toml ADDED Viewed

	@@ -0,0 +1,3 @@

+[theme]
+base="light"
+primaryColor="#29B4E8"

Demo.py ADDED Viewed

	@@ -0,0 +1,163 @@

+import streamlit as st
+import sparknlp
+import os
+import pandas as pd
+from sparknlp.base import *
+from sparknlp.annotator import *
+from pyspark.ml import Pipeline
+from sparknlp.pretrained import PretrainedPipeline
+from annotated_text import annotated_text
+from pyspark.sql import SparkSession
+from pyspark.sql.functions import col, concat, lit, round
+# Page configuration
+st.set_page_config(
+    layout="wide",
+    page_title="Spark NLP Demos App",
+    initial_sidebar_state="auto"
+)
+# CSS for styling
+st.markdown("""
+    <style>
+        .main-title {
+            font-size: 36px;
+            color: #4A90E2;
+            font-weight: bold;
+            text-align: center;
+        }
+        .section p, .section ul {
+            color: #666666;
+        }
+    </style>
+""", unsafe_allow_html=True)
+@st.cache_resource
+def init_spark():
+    return sparknlp.start()
+@st.cache_resource
+def create_pipeline(model):
+    document_assembler = DocumentAssembler() \
+        .setInputCol("text") \
+        .setOutputCol("document")
+    tokenizer = Tokenizer() \
+        .setInputCols(["document"]) \
+        .setOutputCol("token")
+    embeddings = WordEmbeddingsModel.pretrained('glove_840B_300', lang='xx') \
+        .setInputCols(["document", "token"]) \
+        .setOutputCol("embeddings")
+    ner_model = NerDLModel.pretrained(model, 'xx') \
+        .setInputCols(["document", "token", "embeddings"]) \
+        .setOutputCol("ner")
+    ner_converter = NerConverter() \
+        .setInputCols(["document", "token", "ner"]) \
+        .setOutputCol("ner_chunk")
+    pipeline = Pipeline(stages=[
+        document_assembler,
+        tokenizer,
+        embeddings,
+        ner_model,
+        ner_converter
+    ])
+    return pipeline
+def fit_data(pipeline, data):
+  empty_df = spark.createDataFrame([['']]).toDF('text')
+  pipeline_model = pipeline.fit(empty_df)
+  model = LightPipeline(pipeline_model)
+  result = model.fullAnnotate(data)
+  return result
+def annotate(data):
+    document, chunks, labels = data["Document"], data["NER Chunk"], data["NER Label"]
+    annotated_words = []
+    for chunk, label in zip(chunks, labels):
+        parts = document.split(chunk, 1)
+        if parts[0]:
+            annotated_words.append(parts[0])
+        annotated_words.append((chunk, label))
+        document = parts[1]
+    if document:
+        annotated_words.append(document)
+    annotated_text(*annotated_words)
+# Set up the page layout
+st.markdown('<div class="main-title">Identificare entità generali nel testo italiano with Spark NLP</div>', unsafe_allow_html=True)
+# Sidebar content
+model = st.sidebar.selectbox(
+    "Choose the pretrained model",
+    ["ner_wikiner_glove_840B_300"],
+    help="For more info about the models visit: https://sparknlp.org/models"
+)
+# Reference notebook link in sidebar
+link = """
+<a href="https://colab.research.google.com/github/JohnSnowLabs/spark-nlp-workshop/blob/master/tutorials/streamlit_notebooks/NER_FR.ipynb">
+    <img src="https://colab.research.google.com/assets/colab-badge.svg" style="zoom: 1.3" alt="Open In Colab"/>
+</a>
+"""
+st.sidebar.markdown('Reference notebook:')
+st.sidebar.markdown(link, unsafe_allow_html=True)
+# Load examples
+examples = [
+    """William Henry Gates III (nato il 28 ottobre 1955) è un magnate d'affari americano, sviluppatore di software, investitore e filantropo. È noto soprattutto come co-fondatore di Microsoft Corporation. Durante la sua carriera in Microsoft, Gates ha ricoperto le posizioni di presidente, amministratore delegato (CEO), presidente e capo architetto del software, pur essendo il principale azionista individuale fino a maggio 2014. È uno dei più noti imprenditori e pionieri del rivoluzione dei microcomputer degli anni '70 e '80. Nato e cresciuto a Seattle, Washington, Gates ha co-fondato Microsoft con l'amico d'infanzia Paul Allen nel 1975, ad Albuquerque, nel New Mexico; divenne la più grande azienda di software per personal computer al mondo. Gates ha guidato l'azienda come presidente e CEO fino a quando non si è dimesso da CEO nel gennaio 2000, ma è rimasto presidente e divenne capo architetto del software. Alla fine degli anni '90, Gates era stato criticato per le sue tattiche commerciali, che erano state considerate anticoncorrenziali. Questa opinione è stata confermata da numerose sentenze giudiziarie. Nel giugno 2006, Gates ha annunciato che sarebbe passato a un ruolo part-time presso Microsoft e un lavoro a tempo pieno presso la Bill & Melinda Gates Foundation, la fondazione di beneficenza privata che lui e sua moglie, Melinda Gates, hanno fondato nel 2000. [ 9] A poco a poco trasferì i suoi doveri a Ray Ozzie e Craig Mundie. Si è dimesso da presidente di Microsoft nel febbraio 2014 e ha assunto un nuovo incarico come consulente tecnologico per supportare il neo nominato CEO Satya Nadella.""",
+    """La Gioconda è un dipinto ad olio del XVI secolo creato da Leonardo. Si tiene al Louvre di Parigi.""",
+    """Quando Sebastian Thrun ha iniziato a lavorare su auto a guida autonoma presso Google nel 2007, poche persone al di fuori dell'azienda lo hanno preso sul serio. "Posso dirti che amministratori delegati molto importanti delle principali case automobilistiche americane mi stringerebbero la mano e si allontanerebbero perché non valeva la pena parlarne", ha dichiarato Thrun, ora co-fondatore e CEO della startup di istruzione superiore online Udacity, in un'intervista con Recode all'inizio di questa settimana.""",
+    """Facebook è un servizio di social network lanciato come TheFacebook il 4 febbraio 2004. È stato fondato da Mark Zuckerberg con i suoi compagni di stanza del college e gli altri studenti dell'Università di Harvard Eduardo Saverin, Andrew McCollum, Dustin Moskovitz e Chris Hughes. L'adesione al sito web è stata inizialmente limitata dai fondatori agli studenti di Harvard, ma è stata estesa ad altri college nell'area di Boston, la Ivy League e gradualmente la maggior parte delle università negli Stati Uniti e in Canada.""",
+    """La storia dell'elaborazione del linguaggio naturale iniziò generalmente negli anni '50, sebbene si possano trovare lavori di epoche precedenti. Nel 1950, Alan Turing pubblicò un articolo intitolato "Computing Machinery and Intelligence" che proponeva quello che ora viene chiamato il test di Turing come criterio di intelligenza""",
+    """Geoffrey Everest Hinton è uno psicologo cognitivo e uno scienziato informatico canadese inglese, noto soprattutto per il suo lavoro sulle reti neurali artificiali. Dal 2013 divide il suo tempo lavorando per Google e l'Università di Toronto. Nel 2017 è stato cofondatore ed è diventato Chief Scientific Advisor del Vector Institute di Toronto.""",
+    """Quando ho detto a John che volevo trasferirmi in Alaska, mi ha avvertito che avrei avuto difficoltà a trovare uno Starbucks lì.""",
+    """Steven Paul Jobs era un magnate degli affari americano, designer industriale, investitore e proprietario dei media. È stato presidente, amministratore delegato (CEO) e co-fondatore di Apple Inc., presidente e azionista di maggioranza di Pixar, membro del consiglio di amministrazione di The Walt Disney Company a seguito dell'acquisizione di Pixar, e fondatore, presidente e CEO di NeXT. Jobs è ampiamente riconosciuto come un pioniere della rivoluzione del personal computer degli anni '70 e '80, insieme al co-fondatore di Apple Steve Wozniak. Jobs è nato a San Francisco, in California, e è stato adottato. È cresciuto nella Bay Area di San Francisco. Ha frequentato il Reed College nel 1972 prima di abbandonare quello stesso anno, e ha viaggiato attraverso l'India nel 1974 in cerca di illuminazione e studiando il buddismo Zen.""",
+    """Titanic è un romanzo epico americano del 1997 e film catastrofico diretto, scritto, coprodotto e coprodotto da James Cameron. Incorporando aspetti sia storici che di fantasia, si basa sui racconti dell'affondamento del Titanic RMS e vede protagonisti Leonardo DiCaprio e Kate Winslet come membri di diverse classi sociali che si innamorano a bordo della nave durante il suo viaggio inaugurale sfortunato.""",
+    """Oltre ad essere il re del nord, John Snow è un medico inglese e un leader nello sviluppo dell'anestesia e dell'igiene medica. È considerato il primo a utilizzare i dati per curare l'epidemia di colera nel 1834."""
+]
+# st.subheader("Riconoscere persone, luoghi, organizzazioni e varie entità utilizzando un modello di apprendimento profondo preconfigurato predefinito e incorporamenti di parole GloVe (glove_300d).")
+selected_text = st.selectbox("Select an example", examples)
+custom_input = st.text_input("Try it with your own Sentence!")
+text_to_analyze = custom_input if custom_input else selected_text
+st.subheader('Full example text')
+HTML_WRAPPER = """<div class="scroll entities" style="overflow-x: auto; border: 1px solid #e6e9ef; border-radius: 0.25rem; padding: 1rem; margin-bottom: 2.5rem; white-space:pre-wrap">{}</div>"""
+st.markdown(HTML_WRAPPER.format(text_to_analyze), unsafe_allow_html=True)
+# Initialize Spark and create pipeline
+spark = init_spark()
+pipeline = create_pipeline(model)
+output = fit_data(pipeline, text_to_analyze)
+# Display matched sentence
+st.subheader("Processed output:")
+results = {
+    'Document': output[0]['document'][0].result,
+    'NER Chunk': [n.result for n in output[0]['ner_chunk']],
+    'NER Label': [n.metadata['entity'] for n in output[0]['ner_chunk']],
+    'Confidence': [n.metadata['confidence'] for n in output[0]['ner_chunk']]
+}
+spark_df = spark.createDataFrame(
+    zip(results['NER Chunk'], results['NER Label'], results['Confidence']),
+    schema=['NER Chunk', 'NER Label', 'Confidence']
+)
+processed_df = spark_df.select(
+    col('NER Chunk').alias('result'),
+    col('NER Label').alias('entity'),
+    concat(round(col('Confidence').cast('float') * 100, 2), lit('%')).alias('confidence')
+)
+annotate(results)
+with st.expander("View DataFrame"):
+    st.dataframe(processed_df)

Dockerfile ADDED Viewed

	@@ -0,0 +1,70 @@

+# Download base image ubuntu 18.04
+FROM ubuntu:18.04
+# Set environment variables
+ENV NB_USER jovyan
+ENV NB_UID 1000
+ENV HOME /home/${NB_USER}
+# Install required packages
+RUN apt-get update && apt-get install -y \
+    tar \
+    wget \
+    bash \
+    rsync \
+    gcc \
+    libfreetype6-dev \
+    libhdf5-serial-dev \
+    libpng-dev \
+    libzmq3-dev \
+    python3 \
+    python3-dev \
+    python3-pip \
+    unzip \
+    pkg-config \
+    software-properties-common \
+    graphviz \
+    openjdk-8-jdk \
+    ant \
+    ca-certificates-java \
+    && apt-get clean \
+    && update-ca-certificates -f;
+# Install Python 3.8 and pip
+RUN add-apt-repository ppa:deadsnakes/ppa \
+    && apt-get update \
+    && apt-get install -y python3.8 python3-pip \
+    && apt-get clean;
+# Set up JAVA_HOME
+ENV JAVA_HOME /usr/lib/jvm/java-8-openjdk-amd64/
+RUN mkdir -p ${HOME} \
+    && echo "export JAVA_HOME=/usr/lib/jvm/java-8-openjdk-amd64/" >> ${HOME}/.bashrc \
+    && chown -R ${NB_UID}:${NB_UID} ${HOME}
+# Create a new user named "jovyan" with user ID 1000
+RUN useradd -m -u ${NB_UID} ${NB_USER}
+# Switch to the "jovyan" user
+USER ${NB_USER}
+# Set home and path variables for the user
+ENV HOME=/home/${NB_USER} \
+    PATH=/home/${NB_USER}/.local/bin:$PATH
+# Set the working directory to the user's home directory
+WORKDIR ${HOME}
+# Upgrade pip and install Python dependencies
+RUN python3.8 -m pip install --upgrade pip
+COPY requirements.txt /tmp/requirements.txt
+RUN python3.8 -m pip install -r /tmp/requirements.txt
+# Copy the application code into the container at /home/jovyan
+COPY --chown=${NB_USER}:${NB_USER} . ${HOME}
+# Expose port for Streamlit
+EXPOSE 7860
+# Define the entry point for the container
+ENTRYPOINT ["streamlit", "run", "Demo.py", "--server.port=7860", "--server.address=0.0.0.0"]

pages/Workflow & Model Overview.py ADDED Viewed

	@@ -0,0 +1,242 @@

+import streamlit as st
+# Custom CSS for better styling
+st.markdown("""
+    <style>
+        .main-title {
+            font-size: 36px;
+            color: #4A90E2;
+            font-weight: bold;
+            text-align: center;
+        }
+        .sub-title {
+            font-size: 24px;
+            color: #4A90E2;
+            margin-top: 20px;
+        }
+        .section {
+            background-color: #f9f9f9;
+            padding: 15px;
+            border-radius: 10px;
+            margin-top: 20px;
+        }
+        .section h2 {
+            font-size: 22px;
+            color: #4A90E2;
+        }
+        .section p, .section ul {
+            color: #666666;
+        }
+        .link {
+            color: #4A90E2;
+            text-decoration: none;
+        }
+    </style>
+""", unsafe_allow_html=True)
+# Main Title
+st.markdown('<div class="main-title">State-of-the-Art Named Entity Recognition with Spark NLP (Italian)</div>', unsafe_allow_html=True)
+# Introduction
+st.markdown("""
+<div class="section">
+    <p>Named Entity Recognition (NER) is the task of identifying important words in a text and associating them with a category. For example, we may be interested in finding all the personal names in documents, or company names in news articles. Other examples include domain-specific uses such as identifying all disease names in a clinical text, or company trading codes in financial ones.</p>
+    <p>NER can be implemented with many approaches. In this post, we introduce a deep learning-based method using the NerDL model. This approach leverages the scalability of Spark NLP with Python.</p>
+</div>
+""", unsafe_allow_html=True)
+# Introduction to Spark NLP
+st.markdown('<div class="sub-title">Introduction to Spark NLP</div>', unsafe_allow_html=True)
+st.markdown("""
+<div class="section">
+    <p>Spark NLP is an open-source library maintained by John Snow Labs. It is built on top of Apache Spark and Spark ML and provides simple, performant & accurate NLP annotations for machine learning pipelines that can scale easily in a distributed environment.</p>
+    <p>To install Spark NLP, you can simply use any package manager like conda or pip. For example, using pip you can simply run <code>pip install spark-nlp</code>. For different installation options, check the official <a href="https://nlp.johnsnowlabs.com/docs/en/install" target="_blank" class="link">documentation</a>.</p>
+</div>
+""", unsafe_allow_html=True)
+# Using NerDL Model
+st.markdown('<div class="sub-title">Using NerDL Model</div>', unsafe_allow_html=True)
+st.markdown("""
+<div class="section">
+    <p>The NerDL model in Spark NLP is a deep learning-based approach for NER tasks. It uses a Char CNNs - BiLSTM - CRF architecture that achieves state-of-the-art results in most datasets. The training data should be a labeled Spark DataFrame in the format of CoNLL 2003 IOB with annotation type columns.</p>
+</div>
+""", unsafe_allow_html=True)
+# Setup Instructions
+st.markdown('<div class="sub-title">Setup</div>', unsafe_allow_html=True)
+st.markdown('<p>To install Spark NLP in Python, use your favorite package manager (conda, pip, etc.). For example:</p>', unsafe_allow_html=True)
+st.code("""
+pip install spark-nlp
+pip install pyspark
+""", language="bash")
+st.markdown("<p>Then, import Spark NLP and start a Spark session:</p>", unsafe_allow_html=True)
+st.code("""
+import sparknlp
+# Start Spark Session
+spark = sparknlp.start()
+""", language='python')
+# Example Usage with NerDL Model in Italian
+st.markdown('<div class="sub-title">Example Usage with NerDL Model in Italian</div>', unsafe_allow_html=True)
+st.markdown("""
+<div class="section">
+    <p>Below is an example of how to set up and use the NerDL model for named entity recognition in Italian:</p>
+</div>
+""", unsafe_allow_html=True)
+st.code('''
+from sparknlp.base import *
+from sparknlp.annotator import *
+from pyspark.ml import Pipeline
+from pyspark.sql.functions import col, expr, round, concat, lit
+# Document Assembler
+document_assembler = DocumentAssembler() \\
+    .setInputCol("text") \\
+    .setOutputCol("document")
+# Tokenizer
+tokenizer = Tokenizer() \\
+    .setInputCols(["document"]) \\
+    .setOutputCol("token")
+# Word Embeddings
+embeddings = WordEmbeddingsModel.pretrained('glove_840B_300', lang='xx') \\
+    .setInputCols(["document", "token"]) \\
+    .setOutputCol("embeddings")
+# NerDL Model
+ner_model = NerDLModel.pretrained('ner_wikiner_glove_840B_300', 'xx') \\
+    .setInputCols(["document", "token", "embeddings"]) \\
+    .setOutputCol("ner")
+# NER Converter
+ner_converter = NerConverter() \\
+    .setInputCols(["document", "token", "ner"]) \\
+    .setOutputCol("ner_chunk")
+# Pipeline
+pipeline = Pipeline(stages=[
+    document_assembler,
+    tokenizer,
+    embeddings,
+    ner_model,
+    ner_converter
+])
+# Example sentence
+example = """
+Giuseppe Verdi nacque a Le Roncole, una frazione di Busseto, il 10 ottobre 1813.
+Era un compositore italiano, noto per le sue opere come La Traviata e Rigoletto.
+"""
+data = spark.createDataFrame([[example]]).toDF("text")
+# Transforming data
+result = pipeline.fit(data).transform(data)
+# Select the result, entity, and confidence columns
+result.select(
+    expr("explode(ner_chunk) as ner_chunk")
+).select(
+    col("ner_chunk.result").alias("result"),
+    col("ner_chunk.metadata").getItem("entity").alias("entity"),
+    concat(
+        round((col("ner_chunk.metadata").getItem("confidence").cast("float") * 100), 2),
+        lit("%")
+    ).alias("confidence")
+).show(truncate=False)
+''', language="python")
+st.text("""
++--------------+------+----------+
+|result        |entity|confidence|
++--------------+------+----------+
+|Giuseppe Verdi|PER   |66.19%    |
+|Le Roncole    |LOC   |51.45%    |
+|Busseto       |LOC   |85.39%    |
+|La Traviata   |MISC  |69.35%    |
+|Rigoletto     |MISC  |93.34%    |
++--------------+------+----------+
+""")
+# Benchmark Section
+st.markdown('<div class="sub-title">Benchmark</div>', unsafe_allow_html=True)
+st.markdown("""
+<div class="section">
+    <p>Evaluating the performance of NER models is crucial to understanding their effectiveness in real-world applications. Below are the benchmark results for the "ner_wikiner_glove_840B_300" model on Italian text, focusing on various named entity categories. The metrics used include precision, recall, and F1-score, which are standard for evaluating classification models.</p>
+</div>
+""", unsafe_allow_html=True)
+st.markdown("""
+<div class="sub-title">Detailed Results</div>
+| Entity Type | Precision | Recall | F1-Score | Support |
+|-------------|:---------:|:------:|:--------:|--------:|
+| **B-LOC**   | 0.88      | 0.92   | 0.90     | 13,050  |
+| **I-ORG**   | 0.78      | 0.71   | 0.74     | 1,211   |
+| **I-LOC**   | 0.89      | 0.85   | 0.87     | 7,454   |
+| **I-PER**   | 0.93      | 0.94   | 0.94     | 4,539   |
+| **B-ORG**   | 0.88      | 0.72   | 0.79     | 2,222   |
+| **B-PER**   | 0.90      | 0.93   | 0.92     | 7,206   |
+<div class="sub-title">Averages</div>
+| Average Type | Precision | Recall | F1-Score | Support |
+|--------------|:---------:|:------:|:--------:|--------:|
+| **Micro**    | 0.89      | 0.89   | 0.89     | 35,682  |
+| **Macro**    | 0.88      | 0.85   | 0.86     | 35,682  |
+| **Weighted** | 0.89      | 0.89   | 0.89     | 35,682  |
+<div class="sub-title">Category-Specific Performance</div>
+| Category | Precision | Recall | F1-Score | Support |
+|----------|:---------:|:------:|:--------:|--------:|
+| **LOC**  | 86.33%    | 90.53% | 88.38    | 13,685  |
+| **MISC** | 81.88%    | 67.03% | 73.72    | 3,069   |
+| **ORG**  | 85.91%    | 70.52% | 77.46    | 1,824   |
+| **PER**  | 89.54%    | 92.08% | 90.79    | 7,410   |
+<div class="sub-title">Additional Metrics</div>
+- **Processed Tokens:** 349,242
+- **Total Phrases:** 26,227
+- **Found Phrases:** 25,988
+- **Correct Phrases:** 22,529
+- **Accuracy (non-O):** 85.99%
+- **Overall Accuracy:** 98.06%
+""", unsafe_allow_html=True)
+# Summary
+st.markdown('<div class="sub-title">Summary</div>', unsafe_allow_html=True)
+st.markdown("""
+<div class="section">
+    <p>In this article, we discussed named entity recognition using a deep learning-based method with the "wikiner_840B_300" model for Italian. We introduced how to perform the task using the open-source Spark NLP library with Python, which can be used at scale in the Spark ecosystem. These methods can be used for natural language processing applications in various fields, including finance and healthcare.</p>
+</div>
+""", unsafe_allow_html=True)
+# References
+st.markdown('<div class="sub-title">References</div>', unsafe_allow_html=True)
+st.markdown("""
+<div class="section">
+    <ul>
+        <li><a class="link" href="https://sparknlp.org/api/python/reference/autosummary/sparknlp/annotator/ner/ner_dl/index.html" target="_blank" rel="noopener">NerDLModel</a> annotator documentation</li>
+        <li>Model Used: <a class="link" href="https://sparknlp.org/2021/07/19/ner_wikiner_glove_840B_300_xx.html" target="_blank" rel="noopener">ner_wikiner_glove_840B_300</a></li>
+        <li><a class="link" href="https://nlp.johnsnowlabs.com/recognize_entitie" target="_blank" rel="noopener">Visualization demos for NER in Spark NLP</a></li>
+        <li><a class="link" href="https://www.johnsnowlabs.com/named-entity-recognition-ner-with-bert-in-spark-nlp/">Named Entity Recognition (NER) with BERT in Spark NLP</a></li>
+    </ul>
+</div>
+""", unsafe_allow_html=True)
+# Community & Support
+st.markdown('<div class="sub-title">Community & Support</div>', unsafe_allow_html=True)
+st.markdown("""
+<div class="section">
+    <ul>
+        <li><a class="link" href="https://sparknlp.org/" target="_blank">Official Website</a>: Documentation and examples</li>
+        <li><a class="link" href="https://join.slack.com/t/spark-nlp/shared_invite/zt-198dipu77-L3UWNe_AJf4Rqb3DaMb-7A" target="_blank">Slack Community</a>: Connect with other Spark NLP users</li>
+        <li><a class="link" href="https://github.com/JohnSnowLabs/spark-nlp" target="_blank">GitHub Repository</a>: Source code and issue tracker</li>
+    </ul>
+</div>
+""", unsafe_allow_html=True)

requirements.txt ADDED Viewed

	@@ -0,0 +1,6 @@

+streamlit
+st-annotated-text
+pandas
+numpy
+spark-nlp
+pyspark