bbstrader.models¶
Quant toolkit for signals and risk: NLP sentiment analysis / topic modeling, and portfolio optimization.
models ¶
Overview¶
The Models Module provides a collection of quantitative models for financial analysis and decision-making. It includes tools for portfolio optimization and natural language processing (NLP) to extract insights from financial text data. This module is designed to support quantitative trading strategies by providing a robust framework for financial modeling.
Features¶
- Portfolio Optimization: Implements techniques to optimize portfolio allocation, helping to maximize returns and manage risk.
- Natural Language Processing (NLP): Provides tools for analyzing financial news and other text-based data to gauge market sentiment.
- Extensible Design: Structured to allow for the easy addition of new quantitative models and algorithms.
Components¶
- Optimization: Contains portfolio optimization models and related utilities.
- NLP: Includes tools and models for natural language processing tailored for financial applications.
Examples¶
from bbstrader.models import optimized_weights
Assuming 'returns' is a DataFrame of asset returns¶
optimal_weights = optimized_weights(returns=returns) print(optimal_weights)
Notes¶
This module is focused on providing the analytical tools for quantitative analysis. The models can be integrated into trading strategies to provide data-driven signals.
SentimentAnalyzer ¶
Bases: object
A financial sentiment analysis tool that processes and analyzes sentiment from news articles, social media posts, and financial reports.
This class utilizes NLP techniques to preprocess text and apply sentiment analysis using VADER (SentimentIntensityAnalyzer) and optional TextBlob for enhanced polarity scoring.
Initializes the SentimentAnalyzer class by downloading necessary NLTK resources and loading the SpaCy NLP model.
- Downloads NLTK tokenization (
punkt) and stopwords. - Loads the
en_core_web_smSpaCy model with Named Entity Recognition (NER) disabled. - Initializes VADER's SentimentIntensityAnalyzer for sentiment scoring.
Source code in src/bbstrader/models/nlp.py
preprocess_text ¶
Preprocesses the input text by performing the following steps: 1. Converts text to lowercase. 2. Removes URLs. 3. Removes all non-alphabetic characters (punctuation, numbers, special symbols). 4. Tokenizes the text into words. 5. Removes stop words. 6. Lemmatizes the words using SpaCy, excluding pronouns.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
The input text to preprocess. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
str |
The cleaned and lemmatized text. |
Source code in src/bbstrader/models/nlp.py
analyze_sentiment ¶
Analyzes the sentiment of a list of texts using VADER or TextBlob.
Steps:
1. If a custom lexicon is provided, updates the VADER lexicon.
2. If textblob is set to True, computes sentiment using TextBlob.
3. Otherwise, preprocesses the text and computes sentiment using VADER.
4. Returns the average sentiment score of all input texts.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
texts
|
list of str
|
A list of text inputs to analyze. |
required |
lexicon
|
dict
|
A custom sentiment lexicon to update VADER's default lexicon. |
None
|
textblob
|
bool
|
If True, uses TextBlob for sentiment analysis instead of VADER. |
False
|
Returns:
| Name | Type | Description |
|---|---|---|
float |
float
|
The average sentiment score across all input texts. - Positive values indicate positive sentiment. - Negative values indicate negative sentiment. - Zero indicates neutral sentiment. |
Source code in src/bbstrader/models/nlp.py
get_sentiment_for_tickers ¶
get_sentiment_for_tickers(tickers: List[str] | List[Tuple[str, str]], lexicon=None, asset_type='stock', top_news=10, **kwargs) -> Dict[str, float]
Compute sentiment scores for a list of financial tickers based on news and social media data.
Process¶
- Collect news articles and posts related to each ticker from various sources:
- Yahoo Finance News
- Google Finance News
- Reddit posts
- Financial Modeling Prep (FMP) news
- Analyze sentiment from each source:
- Uses VADER for Yahoo and Google Finance news.
- Uses TextBlob for Reddit and FMP news.
- Compute an overall sentiment score using a weighted average approach.
Parameters¶
tickers : list of str or list of tuple A list of asset tickers to analyze. * If using tuples, the first element is the ticker and the second is the asset type. * If using a single string, the asset type must be specified or defaults to "stock". lexicon : dict, optional A custom sentiment lexicon to update VADER's default lexicon. Default is None. asset_type : str, optional The type of asset. Default is "stock". Supported types include: * "stock": Stock symbols (e.g., AAPL, MSFT) * "etf": Exchange-traded funds (e.g., SPY, QQQ) * "future": Futures contracts (e.g., CL=F for crude oil) * "forex": Forex pairs (e.g., EURUSD=X, USDJPY=X) * "crypto": Cryptocurrency pairs (e.g., BTC-USD, ETH-USD) * "index": Stock market indices (e.g., ^GSPC for S&P 500) top_news : int, optional Number of news articles/posts to fetch per source. Default is 10. **kwargs : dict Additional parameters for API authentication and data retrieval. Must include: * fmp_api (str): API key for Financial Modeling Prep. * client_id, client_secret, user_agent (str): Credentials for Reddit API.
Returns¶
dict of str to float A dictionary mapping each ticker to its overall sentiment score. * Positive values indicate positive sentiment. * Negative values indicate negative sentiment. * Zero indicates neutral sentiment.
Notes¶
Ticker names must follow Yahoo Finance conventions.
Source code in src/bbstrader/models/nlp.py
589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 | |
get_topn_sentiments ¶
Retrieves the top and bottom N assets based on sentiment scores.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
sentiments
|
dict
|
A dictionary mapping asset tickers to their sentiment scores. |
required |
topn
|
int
|
The number of top and bottom assets to return. Defaults to 10. |
10
|
Returns:
| Name | Type | Description |
|---|---|---|
tuple |
A tuple containing two lists:
- bottom (list of tuples): The |
Source code in src/bbstrader/models/nlp.py
visualize_sentiments ¶
Visualizes sentiment scores for financial assets using different chart types.
Visualization Modes: - "bar": Displays a bar chart of the top N assets by sentiment score. - "scatter": Displays a scatter plot of sentiment scores.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
sentiment_dict
|
dict
|
A dictionary mapping asset tickers to their sentiment scores. |
required |
mode
|
str
|
The type of visualization to generate. Options: "bar" (default), "scatter". |
'bar'
|
top_n
|
int
|
The number of top tickers to display in the bar chart. Only applicable when mode is "bar". |
10
|
Returns:
| Name | Type | Description |
|---|---|---|
None |
Displays the sentiment visualization. |
Source code in src/bbstrader/models/nlp.py
markowitz_weights ¶
Calculates optimal portfolio weights using Markowitz's mean-variance optimization (Max Sharpe Ratio or Min Volatility) with multiple solvers.
Parameters¶
prices : pd.DataFrame, optional Price data for assets, where rows represent time periods and columns represent assets. rfr : float, optional Risk-free rate (default is 0.0). freq : int, optional Frequency of the data, such as 252 for daily returns in a year (default is 252). min_vol : bool, optional If True, optimizes for minimum volatility instead of maximum Sharpe ratio (default is False).
Returns¶
dict Dictionary containing the optimal asset weights for maximizing the Sharpe ratio or minimizing volatility, normalized to sum to 1.
Notes¶
This function attempts to maximize the Sharpe ratio by iterating through various solvers ('SCS', 'ECOS', 'OSQP') from the PyPortfolioOpt library. If a solver fails, it proceeds to the next one. If none succeed, an error message is printed for each solver that fails.
This function is useful for portfolio with a small number of assets, as it may not scale well for large portfolios.
Raises¶
Exception If all solvers fail, each will print an exception error message during runtime.
Source code in src/bbstrader/models/optimization.py
hierarchical_risk_parity ¶
Computes asset weights using Hierarchical Risk Parity (HRP) for risk-averse portfolio allocation.
Parameters¶
prices : pd.DataFrame, optional
Price data for assets; if provided, daily returns will be calculated.
returns : pd.DataFrame, optional
Daily returns for assets. One of prices or returns must be provided.
freq : int, optional
Number of days to consider in calculating portfolio weights (default is 252).
Returns¶
dict Optimized asset weights using the HRP method, with asset weights summing to 1.
Raises¶
ValueError
If neither prices nor returns are provided.
Notes¶
Hierarchical Risk Parity is particularly useful for portfolios with a large number of assets, as it mitigates issues of multicollinearity and estimation errors in covariance matrices by using hierarchical clustering.
Source code in src/bbstrader/models/optimization.py
equal_weighted ¶
Generates an equal-weighted portfolio by assigning an equal proportion to each asset.
Parameters¶
prices : pd.DataFrame, optional
Price data for assets, where each column represents an asset.
returns : pd.DataFrame, optional
Return data for assets. One of prices or returns must be provided.
round_digits : int, optional
Number of decimal places to round each weight to (default is 5).
Returns¶
dict Dictionary with equal weights assigned to each asset, summing to 1.
Raises¶
ValueError
If neither prices nor returns are provided.
Notes¶
Equal weighting is a simple allocation method that assumes equal importance across all assets, useful as a baseline model and when no strong views exist on asset return expectations or risk.
Source code in src/bbstrader/models/optimization.py
black_litterman_weights ¶
black_litterman_weights(prices=None, rfr=0.0, freq=252, views=None, view_confidences=None, pi=None, market_caps=None)
Computes portfolio weights using the Black-Litterman model.
Parameters¶
prices : pd.DataFrame Price data for assets. rfr : float, optional Risk-free rate (default is 0.0). freq : int, optional Frequency of the data (default is 252). views : dict, optional Investor's views on asset returns. view_confidences : list or np.array, optional Confidence levels for each view. pi : pd.Series, optional Market-implied prior returns. market_caps : pd.Series, optional Market capitalization of assets.
Returns¶
dict Optimal asset weights based on the Black-Litterman model.
Source code in src/bbstrader/models/optimization.py
optimized_weights ¶
Selects an optimization method to calculate portfolio weights based on user preference.
Parameters¶
prices : pd.DataFrame, optional Price data for assets, required for certain methods. returns : pd.DataFrame, optional Returns data for assets, an alternative input for certain methods. freq : int, optional Number of days for calculating portfolio weights, such as 252 for a year's worth of daily returns (default is 252). method : str, optional Optimization method to use ('markowitz', 'hrp', or 'equal') (default is 'equal').
Returns¶
dict Dictionary containing optimized asset weights based on the chosen method.
Raises¶
ValueError If an unknown optimization method is specified.
Notes¶
This function integrates different optimization methods: - 'markowitz': mean-variance optimization with max Sharpe ratio - 'hrp': Hierarchical Risk Parity, for risk-based clustering of assets - 'equal': Equal weighting across all assets
Source code in src/bbstrader/models/optimization.py
nlp ¶
SentimentAnalyzer ¶
Bases: object
A financial sentiment analysis tool that processes and analyzes sentiment from news articles, social media posts, and financial reports.
This class utilizes NLP techniques to preprocess text and apply sentiment analysis using VADER (SentimentIntensityAnalyzer) and optional TextBlob for enhanced polarity scoring.
Initializes the SentimentAnalyzer class by downloading necessary NLTK resources and loading the SpaCy NLP model.
- Downloads NLTK tokenization (
punkt) and stopwords. - Loads the
en_core_web_smSpaCy model with Named Entity Recognition (NER) disabled. - Initializes VADER's SentimentIntensityAnalyzer for sentiment scoring.
Source code in src/bbstrader/models/nlp.py
preprocess_text ¶
Preprocesses the input text by performing the following steps: 1. Converts text to lowercase. 2. Removes URLs. 3. Removes all non-alphabetic characters (punctuation, numbers, special symbols). 4. Tokenizes the text into words. 5. Removes stop words. 6. Lemmatizes the words using SpaCy, excluding pronouns.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
The input text to preprocess. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
str |
The cleaned and lemmatized text. |
Source code in src/bbstrader/models/nlp.py
analyze_sentiment ¶
Analyzes the sentiment of a list of texts using VADER or TextBlob.
Steps:
1. If a custom lexicon is provided, updates the VADER lexicon.
2. If textblob is set to True, computes sentiment using TextBlob.
3. Otherwise, preprocesses the text and computes sentiment using VADER.
4. Returns the average sentiment score of all input texts.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
texts
|
list of str
|
A list of text inputs to analyze. |
required |
lexicon
|
dict
|
A custom sentiment lexicon to update VADER's default lexicon. |
None
|
textblob
|
bool
|
If True, uses TextBlob for sentiment analysis instead of VADER. |
False
|
Returns:
| Name | Type | Description |
|---|---|---|
float |
float
|
The average sentiment score across all input texts. - Positive values indicate positive sentiment. - Negative values indicate negative sentiment. - Zero indicates neutral sentiment. |
Source code in src/bbstrader/models/nlp.py
get_sentiment_for_tickers ¶
get_sentiment_for_tickers(tickers: List[str] | List[Tuple[str, str]], lexicon=None, asset_type='stock', top_news=10, **kwargs) -> Dict[str, float]
Compute sentiment scores for a list of financial tickers based on news and social media data.
Process¶
- Collect news articles and posts related to each ticker from various sources:
- Yahoo Finance News
- Google Finance News
- Reddit posts
- Financial Modeling Prep (FMP) news
- Analyze sentiment from each source:
- Uses VADER for Yahoo and Google Finance news.
- Uses TextBlob for Reddit and FMP news.
- Compute an overall sentiment score using a weighted average approach.
Parameters¶
tickers : list of str or list of tuple A list of asset tickers to analyze. * If using tuples, the first element is the ticker and the second is the asset type. * If using a single string, the asset type must be specified or defaults to "stock". lexicon : dict, optional A custom sentiment lexicon to update VADER's default lexicon. Default is None. asset_type : str, optional The type of asset. Default is "stock". Supported types include: * "stock": Stock symbols (e.g., AAPL, MSFT) * "etf": Exchange-traded funds (e.g., SPY, QQQ) * "future": Futures contracts (e.g., CL=F for crude oil) * "forex": Forex pairs (e.g., EURUSD=X, USDJPY=X) * "crypto": Cryptocurrency pairs (e.g., BTC-USD, ETH-USD) * "index": Stock market indices (e.g., ^GSPC for S&P 500) top_news : int, optional Number of news articles/posts to fetch per source. Default is 10. **kwargs : dict Additional parameters for API authentication and data retrieval. Must include: * fmp_api (str): API key for Financial Modeling Prep. * client_id, client_secret, user_agent (str): Credentials for Reddit API.
Returns¶
dict of str to float A dictionary mapping each ticker to its overall sentiment score. * Positive values indicate positive sentiment. * Negative values indicate negative sentiment. * Zero indicates neutral sentiment.
Notes¶
Ticker names must follow Yahoo Finance conventions.
Source code in src/bbstrader/models/nlp.py
589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 | |
get_topn_sentiments ¶
Retrieves the top and bottom N assets based on sentiment scores.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
sentiments
|
dict
|
A dictionary mapping asset tickers to their sentiment scores. |
required |
topn
|
int
|
The number of top and bottom assets to return. Defaults to 10. |
10
|
Returns:
| Name | Type | Description |
|---|---|---|
tuple |
A tuple containing two lists:
- bottom (list of tuples): The |
Source code in src/bbstrader/models/nlp.py
visualize_sentiments ¶
Visualizes sentiment scores for financial assets using different chart types.
Visualization Modes: - "bar": Displays a bar chart of the top N assets by sentiment score. - "scatter": Displays a scatter plot of sentiment scores.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
sentiment_dict
|
dict
|
A dictionary mapping asset tickers to their sentiment scores. |
required |
mode
|
str
|
The type of visualization to generate. Options: "bar" (default), "scatter". |
'bar'
|
top_n
|
int
|
The number of top tickers to display in the bar chart. Only applicable when mode is "bar". |
10
|
Returns:
| Name | Type | Description |
|---|---|---|
None |
Displays the sentiment visualization. |
Source code in src/bbstrader/models/nlp.py
optimization ¶
markowitz_weights ¶
Calculates optimal portfolio weights using Markowitz's mean-variance optimization (Max Sharpe Ratio or Min Volatility) with multiple solvers.
Parameters¶
prices : pd.DataFrame, optional Price data for assets, where rows represent time periods and columns represent assets. rfr : float, optional Risk-free rate (default is 0.0). freq : int, optional Frequency of the data, such as 252 for daily returns in a year (default is 252). min_vol : bool, optional If True, optimizes for minimum volatility instead of maximum Sharpe ratio (default is False).
Returns¶
dict Dictionary containing the optimal asset weights for maximizing the Sharpe ratio or minimizing volatility, normalized to sum to 1.
Notes¶
This function attempts to maximize the Sharpe ratio by iterating through various solvers ('SCS', 'ECOS', 'OSQP') from the PyPortfolioOpt library. If a solver fails, it proceeds to the next one. If none succeed, an error message is printed for each solver that fails.
This function is useful for portfolio with a small number of assets, as it may not scale well for large portfolios.
Raises¶
Exception If all solvers fail, each will print an exception error message during runtime.
Source code in src/bbstrader/models/optimization.py
hierarchical_risk_parity ¶
Computes asset weights using Hierarchical Risk Parity (HRP) for risk-averse portfolio allocation.
Parameters¶
prices : pd.DataFrame, optional
Price data for assets; if provided, daily returns will be calculated.
returns : pd.DataFrame, optional
Daily returns for assets. One of prices or returns must be provided.
freq : int, optional
Number of days to consider in calculating portfolio weights (default is 252).
Returns¶
dict Optimized asset weights using the HRP method, with asset weights summing to 1.
Raises¶
ValueError
If neither prices nor returns are provided.
Notes¶
Hierarchical Risk Parity is particularly useful for portfolios with a large number of assets, as it mitigates issues of multicollinearity and estimation errors in covariance matrices by using hierarchical clustering.
Source code in src/bbstrader/models/optimization.py
equal_weighted ¶
Generates an equal-weighted portfolio by assigning an equal proportion to each asset.
Parameters¶
prices : pd.DataFrame, optional
Price data for assets, where each column represents an asset.
returns : pd.DataFrame, optional
Return data for assets. One of prices or returns must be provided.
round_digits : int, optional
Number of decimal places to round each weight to (default is 5).
Returns¶
dict Dictionary with equal weights assigned to each asset, summing to 1.
Raises¶
ValueError
If neither prices nor returns are provided.
Notes¶
Equal weighting is a simple allocation method that assumes equal importance across all assets, useful as a baseline model and when no strong views exist on asset return expectations or risk.
Source code in src/bbstrader/models/optimization.py
black_litterman_weights ¶
black_litterman_weights(prices=None, rfr=0.0, freq=252, views=None, view_confidences=None, pi=None, market_caps=None)
Computes portfolio weights using the Black-Litterman model.
Parameters¶
prices : pd.DataFrame Price data for assets. rfr : float, optional Risk-free rate (default is 0.0). freq : int, optional Frequency of the data (default is 252). views : dict, optional Investor's views on asset returns. view_confidences : list or np.array, optional Confidence levels for each view. pi : pd.Series, optional Market-implied prior returns. market_caps : pd.Series, optional Market capitalization of assets.
Returns¶
dict Optimal asset weights based on the Black-Litterman model.
Source code in src/bbstrader/models/optimization.py
optimized_weights ¶
Selects an optimization method to calculate portfolio weights based on user preference.
Parameters¶
prices : pd.DataFrame, optional Price data for assets, required for certain methods. returns : pd.DataFrame, optional Returns data for assets, an alternative input for certain methods. freq : int, optional Number of days for calculating portfolio weights, such as 252 for a year's worth of daily returns (default is 252). method : str, optional Optimization method to use ('markowitz', 'hrp', or 'equal') (default is 'equal').
Returns¶
dict Dictionary containing optimized asset weights based on the chosen method.
Raises¶
ValueError If an unknown optimization method is specified.
Notes¶
This function integrates different optimization methods: - 'markowitz': mean-variance optimization with max Sharpe ratio - 'hrp': Hierarchical Risk Parity, for risk-based clustering of assets - 'equal': Equal weighting across all assets