Skip to content

About

A full-stack AI application that uses a locally hosted LLaMA model (via Ollama) to summarize text. This project demonstrates how to integrate a powerful LLM with a Python backend (FastAPI ) and a user-friendly frontend Streamlit

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

🦙 LLaMA Text Summarizer

A full-stack AI application that uses a locally hosted LLaMA model (via Ollama) to summarize text.
Built with a Python backend (FastAPI) and a user-friendly frontend (Streamlit).

LinkedIn Email

Developed as part of Internship Task #1


🚀 Features

  • FastAPI Backend: Acts as the bridge between the user and the AI model.
  • Streamlit Frontend: A clean, interactive web interface for users to input text.
  • Local Inference: Runs entirely offline using Ollama (no API keys required!).
  • Model Agnostic: Can easily switch between llama2, llama3, or lightweight models like phi3.

🧭 How It Works (Flowchart)

```mermaid flowchart TD A[User] -->|Enters text| B[Streamlit Frontend] B -->|Sends HTTP request| C[FastAPI Backend] C -->|Sends prompt| D[Ollama Local Model Runner] D -->|Runs inference| E["LLaMA / Phi3 Model (Local)"] E -->|Returns summary| D D -->|Returns response| C C -->|Sends JSON response| B B -->|Displays summary| A ```

Flow explanation:

  1. The user enters text into the Streamlit UI.
  2. Streamlit sends the text to the FastAPI backend via an HTTP request.
  3. FastAPI forwards the prompt to Ollama, which runs the selected local LLM.
  4. The model generates a summary, which flows back through FastAPI to Streamlit.
  5. The summarized text is displayed to the user.

🛠️ Tech Stack

  • Python 3.10+
  • Ollama (Model Runner)
  • FastAPI (Backend API)
  • Streamlit (Frontend UI)
  • Uvicorn (ASGI Server)

💻 System Requirements (Important!)

Running local AI models requires decent hardware. Here is what you need to run this smoothly:

Requirement Minimum Recommended
RAM 8GB 16GB
GPU Not required NVIDIA GPU with CUDA
Storage 4GB free space 4GB+ free space

Note: If you have 8GB RAM, close your browser tabs before running the model! Without a GPU, the model will run on your CPU — it works, but text generation will be slower.


⚙️ Installation & Setup

1. Install Ollama

Download and install Ollama from ollama.com.

2. Pull the Model

Open your terminal and download the model. Use phi3 or llama2 depending on your RAM.

```bash ollama pull phi3

OR if you have 16GB+ RAM:

ollama pull llama2 ```

3. Clone the Repository

```bash git clone https://github.com//llama-text-summarizer.git cd llama-text-summarizer ```

4. Install Dependencies

```bash pip install -r requirements.txt ```

5. Start the Backend

```bash uvicorn backend.main:app --reload ```

6. Start the Frontend

```bash streamlit run frontend/app.py ```


📂 Project Structure

``` llama-text-summarizer/ ├── backend/ │ └── main.py # FastAPI app ├── frontend/ │ └── app.py # Streamlit UI ├── requirements.txt └── README.md ```


🔮 Future Scope

  • Add support for file uploads (PDF/DOCX summarization).
  • Add adjustable summary length and tone controls.
  • Deploy as a Docker container for easier setup.
  • Add multi-language summarization support.

🤝 Contributing

Contributions are welcome! If you'd like to contribute, please open an issue or submit a pull request.


Made with ❤️ by Neha Maurya
📧 mauryaneha2006@gmail.com  |  🔗 LinkedIn

About

A full-stack AI application that uses a locally hosted LLaMA model (via Ollama) to summarize text. This project demonstrates how to integrate a powerful LLM with a Python backend (FastAPI ) and a user-friendly frontend Streamlit

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages